LLM Torch Profiler Trace Analysis
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…
$ npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install facebookexperimental/triton tma-illegal-instruction --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/tma-illegal-instruction .claude/skills/tma-illegal-instruction && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tma-illegal-instruction" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tma-illegal-instruction into .claude/skills/tma-illegal-instruction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tma-illegal-instruction", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tma-illegal-instructionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install facebookexperimental/triton tma-illegal-instruction --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/tma-illegal-instruction .agents/skills/tma-illegal-instruction && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tma-illegal-instruction" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tma-illegal-instruction into .agents/skills/tma-illegal-instruction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tma-illegal-instruction", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install facebookexperimental/triton tma-illegal-instruction --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/tma-illegal-instruction .cursor/skills/tma-illegal-instruction && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tma-illegal-instruction" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tma-illegal-instruction into .cursor/skills/tma-illegal-instruction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tma-illegal-instruction", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/facebookexperimental/triton.git --path .claude/skills/tma-illegal-instruction--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install facebookexperimental/triton tma-illegal-instruction --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/tma-illegal-instruction .gemini/skills/tma-illegal-instruction && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tma-illegal-instruction" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tma-illegal-instruction into .gemini/skills/tma-illegal-instruction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tma-illegal-instruction", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install facebookexperimental/triton tma-illegal-instructionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/tma-illegal-instruction .github/skills/tma-illegal-instruction && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tma-illegal-instruction" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tma-illegal-instruction into .github/skills/tma-illegal-instruction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tma-illegal-instruction", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install facebookexperimental/triton tma-illegal-instruction --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/tma-illegal-instruction .opencode/skills/tma-illegal-instruction && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tma-illegal-instruction" agent skill from https://github.com/facebookexperimental/triton/tree/main/.claude/skills/tma-illegal-instruction into .opencode/skills/tma-illegal-instruction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tma-illegal-instruction", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tma-illegal-instructionDiagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…
Tma Illegal Instruction is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store, tl.asyncdescriptorload, async TMA copies) as the source code line. Use when the user reports CUDA error 716, "an illegal instruction was encountered", segfault inside a TMA op, kernel hang followed by an illegal instruction trap, or a crash that only fires on the first or last tile of a launch. Covers the pattern where a TMA…
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering GPU and accelerator computing and Root cause analysis. It works with CUDA. The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6f3dd70. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Tma Illegal Instruction loads about 1.1k tokens when it runs. Until then it costs about 199 tokens; SKILL.md has 550 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from facebookexperimental/triton at commit 6f3dd70, republished under its MIT licence (© facebookexperimental). 550 words, ~1,104 tokens.
.claude/skills/tma-illegal-instruction/SKILL.md (or your agent's skills folder).CUDA reports "an illegal instruction was encountered" (error 716), or the
kernel crashes inside a TMA op, on a Triton kernel that uses TMA descriptors
(TensorDescriptor, tl.make_tensor_descriptor, desc.load(...),
desc.store(...), async TMA copies, etc.).
The crash is likely tile-dependent — appears only at certain grid values. This is likely because the tile out of bounds is entirely past the shape of the TME store.
Walk these in order. Don't skip ahead — the first check is the cheapest and the most often correct.
Find the faoiling TMA p. From the stack trace / sanitizer output / IR
dump, identify which descriptor.load(...) or descriptor.store(...)
crashed. Note the offsets it was called with (e.g.
[pid_m * BM, pid_n * BN]) and the descriptor's declared shape.
Reconstruct the failing tile's starting offset. For the failing
program/iteration, compute the literal integer offsets passed to the TMA
op. For each axis i of the descriptor, ask: is off_i >= shape_i?
If yes, that is the bug. The launcher / tile-mapping logic put a program
in a region that does not exist.
Confirm by debug messaging. Determine either the grid or value
(could be a jagged tensor) information that is causing the failure.
Add a tl.device_print call to the kernel with an if that skips the
operation. NOTE: This is the not a proper solution!
Only after the structural bug is identified, determine whether the right fix is launcher/grid dependent or runtime data dependent. If the latter, identify how this shape can be reached.
The common temptation is to wrap the failing TMA op in
if off_m < M and off_n < N: (or to fall back to tl.load with a mask).
Resist this. It silences the symptom but:
tile_id it computed for the previous tiles is
also suspect.In-kernel masks are fine for genuinely ragged shapes (real K not a multiple of BLOCK_K, etc.), but a TMA illegal instruction is a different signal — it says "the launch contract is wrong", not "this iteration is ragged".
For the failing tile/iteration, the kernel should be able to assert
off_i < shape_i for every TMA op. The verification protocol:
tl.device_assert(off_i < shape_i, "...") calls (or print
the offsets) before the suspected TMA op and re-run with the same shape
that crashed.Removing tl.device_assert after verification is required; the structural fix
is what you ship. The code should NOT introduce a new if statement directly over
just the TMA operation (that is typically wrong).
© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/tma-illegal-instruction of facebookexperimental/triton.
Open the folder on GitHubat commit 6f3dd70
Tma Illegal Instruction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tma Illegal Instruction this skillfacebookexperimental/triton | 201 | — | ~1.1k | Automated safety check: Pass | MIT | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.8k | Automated safety check: Pass | None | |
| Cuda Cpp Kernelvipshop/cache-dit | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Cudasablin39/tilelang-cuda-skills | 145 | — | ~990 | Automated safety check: Pass | None | |
| ONNX Runtime CUDA Attention Patternsmicrosoft/onnxruntime | 22k | — | ~6.5k | Automated safety check: Pass | MIT | |
| Doca Gpunetio Ib Write LatNVIDIA/skills | 3.6k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 |
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
sablin39/tilelang-cuda-skills
Draft, debug, and measure CUDA kernels and host launch workflows.
microsoft/onnxruntime
Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.
NVIDIA/skills
A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under…
mohitmishra786/low-level-dev-skills
CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.
facebookexperimental/triton
Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.
facebookexperimental/triton
Design and run Triton TTGIR debugging ablations using iroverride.
facebookexperimental/triton
Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.
facebookexperimental/triton
Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.
facebookexperimental/triton
Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.
facebookexperimental/triton
Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).
Works with
Categories
Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…. Tma Illegal Instruction is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization.asyncdescriptorload, async TMA copies) as the source code line.
Tma Illegal Instruction fits situations like: the user reports CUDA error 716; an illegal instruction was encountered; segfault inside a TMA op; kernel hang followed by an illegal instruction trap.
Run `npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a claude-code`. Or copy the skill folder (.claude/skills/tma-illegal-instruction in facebookexperimental/triton) into .claude/skills/tma-illegal-instruction in your project. Claude Code loads it when a task matches its description.
Run `npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a codex`. Or copy the skill folder (.claude/skills/tma-illegal-instruction in facebookexperimental/triton) into .agents/skills/tma-illegal-instruction in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill tma-illegal-instruction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tma-illegal-instruction, .gemini/skills/tma-illegal-instruction, .github/skills/tma-illegal-instruction and .opencode/skills/tma-illegal-instruction in your project.
SKILL.md names no scripts, command-line tools or credentials: Tma Illegal Instruction is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Tma Illegal Instruction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Tma Illegal Instruction: LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Cuda Cpp Kernel (vipshop/cache-dit, 1.3k stars), Cuda (sablin39/tilelang-cuda-skills, 145 stars) and ONNX Runtime CUDA Attention Patterns (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 10, 2026.
Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.