Official agent skill

Compute Sanitizer

by facebookexperimental in facebookexperimental/triton

Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

OfficialMITAuto-check passedAI & LLM Engineering

Install Compute Sanitizer

skills CLI
$ npx skills add facebookexperimental/triton --skill compute-sanitizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton compute-sanitizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/compute-sanitizer .claude/skills/compute-sanitizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
compute-sanitizer
GitHub stars
201
Token cost
~1.6k tokens
SKILL.md length
683 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

  • Works in 4 steps: Locate the binary. It is often not on… → Get the reproduce command (a… → Pin a healthy GPU. Run bash… → …
  • A kernel produces wrong results
  • SKILL.md covers The four tools, Setup, Invocation and Triton/TLX specifics, plus 2 more sections
  • Calls python and bash

What it does

Compute Sanitizer is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. Use when a kernel produces wrong results, crashes with an illegal/misaligned access, or is suspected of a shared-memory data race or invalid barrier usage — especially warp-specialized (WS) kernels using mbarriers, named barriers, TMA copies, or MMA accumulators. This is a runtime check: it runs the real kernel via its reproduce command, so it needs a working GPU…

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • A kernel produces wrong results
  • Crashes with an illegal/misaligned access
  • Is suspected of a shared-memory data race
  • Invalid barrier usage — especially warp-specialized (WS) kernels using mbarriers

Example prompts

  • “/compute-sanitizer”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Locate the binary. It is often not on PATH. Prefer
  2. Get the reproduce command (a pytest/python invocation). Use the
  3. Pin a healthy GPU. Run bash third_party/tlx/find_working_gpu.sh, take a
  4. Enable source mapping. Triton emits line info by default; keep it on

What it can do on your machine

Read from SKILL.md and the folder at commit 37301d4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Compute Sanitizer loads about 1.6k tokens when it runs. Until then it costs about 144 tokens; SKILL.md has 683 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 37301d4, republished under its MIT licence (© facebookexperimental). 683 words, ~1,639 tokens.

Download SKILL.mdSave it as .claude/skills/compute-sanitizer/SKILL.md (or your agent's skills folder).
name
compute-sanitizer
description
Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. Use when a kernel produces wrong results, crashes with an illegal/misaligned access, or is suspected of a shared-memory data race or invalid barrier usage — especially warp-specialized (WS) kernels using mbarriers, named barriers, TMA copies, or MMA accumulators. This is a runtime check: it runs the real kernel via its reproduce command, so it needs a working GPU and is 10-100x slower than a normal run.

Compute Sanitizer

compute-sanitizer is NVIDIA's runtime correctness checker. It instruments the real kernel launch and reports memory and synchronization errors that static IR analysis cannot see. Use it to confirm (or rule out) out-of-bounds accesses, uninitialized reads, shared-memory data races, and illegal barrier usage in Triton/TLX kernels.

It complements the static investigations: barrier-visualization reasons about the intended barrier protocol from the IR, while compute-sanitizer observes what actually happens at runtime. When a race or sync error fires, cross-check the two.

The four tools

Select with --tool; memcheck is the default.

  • memcheck — out-of-bounds and misaligned global/local/shared accesses, plus device-side allocation/leak errors. First tool to run; catches the classic TMA-past-end and bad-index bugs.
  • racecheck — hazards on shared memory (RAW / WAR / WAW). Highly relevant to WS producer/consumer kernels that stage data through SMEM buffers. Does NOT check global-memory races.
  • synccheck — illegal use of barriers and synchronization primitives (divergent bar.sync, mismatched arrive/wait counts, invalid cluster/async barrier usage). Directly relevant to mbarrier / named-barrier WS code.
  • initcheck — reads of uninitialized global memory. Useful when output has nondeterministic garbage rather than a hard crash.

Setup

  1. Locate the binary. It is often not on PATH. Prefer $COMPUTE_SANITIZER_BIN if set, else command -v compute-sanitizer, else /usr/local/cuda/bin/compute-sanitizer, else the newest /usr/local/cuda-*/bin/compute-sanitizer.
  2. Get the reproduce command (a pytest/python invocation). Use the smallest failing shape/config — sanitizer overhead makes large grids impractical. Prefer a single test id over a whole file.
  3. Pin a healthy GPU. Run bash third_party/tlx/find_working_gpu.sh, take a WORKING_GPUS index, and prepend CUDA_VISIBLE_DEVICES=<idx>. See the debug-failing-gpu skill. If a run hangs past your timeout, run third_party/tlx/killgpu.sh.
  4. Enable source mapping. Triton emits line info by default; keep it on (do not set TRITON_DISABLE_LINE_INFO=1) so errors map to file:line. Pass --show-backtrace yes for host+device backtraces.

Invocation

General form (the sanitizer wraps the whole command, env vars go before it):

bash
CUDA_VISIBLE_DEVICES=<idx> \
  <compute-sanitizer> --tool memcheck \
    --error-exitcode 1 \
    --show-backtrace yes \
    --log-file <out_dir>/memcheck.log \
    --target-processes all \
    python -m pytest -s -x "<test_id>"

Useful flags:

  • --error-exitcode 1 — return non-zero when any error is found (scripting).
  • --log-file <path> — capture the report (%p expands to PID if needed).
  • --target-processes all — follow child processes (pytest workers, subprocs).
  • --kernel-name-exclude / --kernel-name <regex> — filter to the Triton kernel (names look like _attn_fwd_..., matmul_kernel, etc.) to cut noise.
  • --launch-timeout <s> — bound a single launch.
  • memcheck: --leak-check full, --padding <bytes> (catch off-by-a-few OOB).
  • racecheck: --racecheck-report all (hazards + analysis).
  • --print-limit 0 — do not truncate the error list while triaging.

Run the tools in order — memcheck → racecheck → synccheck → initcheck — each into its own log. Stop early only with a stated reason (e.g. memcheck already found the crashing OOB, or an earlier tool hung).

Show full SKILL.md (273 more words)Show less

Triton/TLX specifics

  • First run JIT-compiles the kernel (slow, unrelated to sanitizer). Let the compile complete; do not mistake compile time for a hang.
  • racecheck is shared-memory only. WS kernels stage operands through SMEM, so a missing producer/consumer barrier shows up here as a WAR/RAW hazard on the SMEM buffer. A global-memory race will NOT be reported.
  • synccheck + mbarriers. Invalid named-barrier / mbarrier / cluster-barrier usage (e.g. wrong arrive count, divergent participation) surfaces here. Pair with barrier-visualization to map the offending barrier to a partition.
  • TMA / async copies. Out-of-bounds TMA tile accesses crash with an illegal instruction at runtime rather than masking — memcheck localizes the launch; also see the tma-illegal-instruction skill for the structural launcher-bug pattern.
  • Possible noise. Some async/TMA/cp.async paths may emit benign warnings or unsupported-feature notes on certain driver/toolkit versions. Note them as caveats; do not treat a warning as a confirmed bug without corroboration.

Interpreting output

  • Each tool ends with ========= ERROR SUMMARY: N errors. 0 errors = clean for that tool. racecheck prints RACECHECK SUMMARY.
  • A real finding includes the access kind (e.g. Invalid global read of size 16), the address, the kernel name, and — with line info — the file:line. Capture the first error; later ones are often cascades.
  • "Clean" is a result, not a failure. Report it plainly: a clean memcheck + racecheck + synccheck materially narrows a wrong-result bug toward compiler/logic rather than memory/sync.

Reporting

Return a compact matrix plus triage:

text
tool       | exit | errors | first finding (kernel @ file:line) | mapped?
memcheck   |  0   |   0    | -                                  | n/a
racecheck  |  1   |   3    | WAR hazard on smem buf @ k.py:142  | yes
synccheck  |  0   |   0    | -                                  | n/a
initcheck  |  0   |   0    | -                                  | n/a

Then state the most likely root cause (OOB vs. shared-memory race vs. uninitialized read vs. illegal sync), the implicated TLX/Triton construct, and the next action (e.g. narrow with a kernel filter, or hand a race/sync hit to barrier-visualization).

Do not run performance benchmarks.

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/compute-sanitizer of facebookexperimental/triton.

Open the folder on GitHubat commit 37301d4

Compare with similar skills

Compute Sanitizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Compute Sanitizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Compute Sanitizer this skillfacebookexperimental/triton201—~1.6kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 24 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    188 GitHub stars~547 tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Debug Failing GPU

    facebookexperimental/triton

    Official

    Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

    201 GitHub stars~709 tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed
  • Kernel Perf Testing

    facebookexperimental/triton

    Official

    Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs.

    201 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Compute Sanitizer

What does Compute Sanitizer do?

Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. Compute Sanitizer is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

When should I use Compute Sanitizer?

Compute Sanitizer fits situations like: A kernel produces wrong results; crashes with an illegal/misaligned access; is suspected of a shared-memory data race; invalid barrier usage — especially warp-specialized (WS) kernels using mbarriers.

How do I install Compute Sanitizer in Claude Code?

Run `npx skills add facebookexperimental/triton --skill compute-sanitizer -a claude-code`. Or copy the skill folder (.claude/skills/compute-sanitizer in facebookexperimental/triton) into .claude/skills/compute-sanitizer in your project. Claude Code loads it when a task matches its description.

How do I install Compute Sanitizer in Codex?

Run `npx skills add facebookexperimental/triton --skill compute-sanitizer -a codex`. Or copy the skill folder (.claude/skills/compute-sanitizer in facebookexperimental/triton) into .agents/skills/compute-sanitizer in your project. Codex loads it when a task matches its description.

Can I use Compute Sanitizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill compute-sanitizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compute-sanitizer, .gemini/skills/compute-sanitizer, .github/skills/compute-sanitizer and .opencode/skills/compute-sanitizer in your project.

What does Compute Sanitizer need to run?

Going by SKILL.md and its folder, Compute Sanitizer needs the command-line tools its instructions call (python and bash). Our summary lists: Python 3.

Does Compute Sanitizer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Compute Sanitizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Compute Sanitizer use?

Compute Sanitizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Compute Sanitizer use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Compute Sanitizer?

Skills that share tags, products or a category with Compute Sanitizer: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Compute Sanitizer?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 11, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.