Official agent skill

Debug Failing GPU

by facebookexperimental in facebookexperimental/triton

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

OfficialMITAuto-check passedAI & LLM Engineering

Install Debug Failing GPU

skills CLI
$ npx skills add facebookexperimental/triton --skill debug-failing-gpu -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install facebookexperimental/triton debug-failing-gpu --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/facebookexperimental/triton.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/debug-failing-gpu .claude/skills/debug-failing-gpu && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-failing-gpu
GitHub stars
201
Token cost
~709 tokens
SKILL.md length
247 words
Files
1
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

  • Works in 4 steps: Run the scanner → Read the final line, WORKING_GPUS=...… → Pick the first working index. If several… → …
  • A command (pytest
  • SKILL.md covers When to trigger, Procedure, If no GPU works and Command-rewrite examples
  • Calls bash and pytest

What it does

Debug Failing GPU is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Recover from GPU-busy / GPU-unavailable failures. Use when a command (pytest, python, a TLX/Triton kernel run, a benchmark) fails with errors indicating the GPU is busy, out of memory, or unavailable — e.g. "CUDA error: out of memory", "all CUDA-capable devices are busy or unavailable", "CUDA-capable device(s) is/are busy or unavailable", "RuntimeError: No CUDA GPUs are available", "device-side assert", or a hang on the first CUDA call. Runs findworkinggpu.sh to locate a healthy GPU and re-runs the failed command…

Its SKILL.md is about 710 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering GPU and accelerator computing and Unit testing. It works with CUDA, pytest and Python. The repository describes itself as: Github mirror of trition-lang/triton repo. The licence is MIT.

When your agent uses it

  • A command (pytest
  • A TLX/Triton kernel run
  • A benchmark) fails with errors indicating the GPU is busy
  • Unavailable — e.g

Example prompts

  • “CUDA error: out of memory”
  • “all CUDA-capable devices are busy or unavailable”
  • “CUDA-capable device(s) is/are busy or unavailable”
  • “/debug-failing-gpu”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run the scanner
  2. Read the final line, WORKING_GPUS=... (e.g. WORKING_GPUS=0,2,3). These are
  3. Pick the first working index. If several are free, mention them so the user can
  4. Re-issue the original failing command with CUDA_VISIBLE_DEVICES=

What it can do on your machine

Read from SKILL.md and the folder at commit 6f3dd70. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Failing GPU loads about 709 tokens when it runs. Until then it costs about 144 tokens; SKILL.md has 247 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~709

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from facebookexperimental/triton at commit 6f3dd70, republished under its MIT licence (© facebookexperimental). 247 words, ~709 tokens.

Download SKILL.mdSave it as .claude/skills/debug-failing-gpu/SKILL.md (or your agent's skills folder).
name
debug-failing-gpu
description
Recover from GPU-busy / GPU-unavailable failures. Use when a command (pytest, python, a TLX/Triton kernel run, a benchmark) fails with errors indicating the GPU is busy, out of memory, or unavailable — e.g. "CUDA error: out of memory", "all CUDA-capable devices are busy or unavailable", "CUDA-capable device(s) is/are busy or unavailable", "RuntimeError: No CUDA GPUs are available", "device-side assert", or a hang on the first CUDA call. Runs find_working_gpu.sh to locate a healthy GPU and re-runs the failed command pinned to it via CUDA_VISIBLE_DEVICES.

Debug Failing GPU

A command failed because the GPU it landed on is busy, out of memory, or in a bad state. Find a GPU that actually works and re-run the command pinned to it.

When to trigger

Any failure whose root cause is the device, not the code. Common signatures:

  • CUDA error: out of memory / torch.cuda.OutOfMemoryError
  • all CUDA-capable devices are busy or unavailable
  • CUDA-capable device(s) is/are busy or unavailable
  • RuntimeError: No CUDA GPUs are available
  • CUDA error: device-side assert triggered
  • A kernel/test that hangs on the first CUDA call

Do not use this for kernel logic bugs, compilation errors, or numerical mismatches — those are not device-health problems.

Procedure

  1. Run the scanner:
    bash
    bash third_party/tlx/find_working_gpu.sh
  2. Read the final line, WORKING_GPUS=... (e.g. WORKING_GPUS=0,2,3). These are physical GPU indices.
  3. Pick the first working index. If several are free, mention them so the user can parallelize across GPUs.
  4. Re-issue the original failing command with CUDA_VISIBLE_DEVICES=<idx>:
    • If the command had no CUDA_VISIBLE_DEVICES, prepend one.
    • If it already set CUDA_VISIBLE_DEVICES, replace that value — do not stack two assignments.

If no GPU works

WORKING_GPUS= is empty. The GPUs may be held by your own stuck processes:

  1. Clear them: third_party/tlx/killgpu.sh
  2. Re-run bash third_party/tlx/find_working_gpu.sh.
  3. If still empty, all GPUs are occupied by other users — report that and stop; there is nothing to switch to.

Command-rewrite examples

bash
# No device set -> prepend
pytest third_party/tlx/tutorials/testing/test_correctness.py
# becomes
CUDA_VISIBLE_DEVICES=2 pytest third_party/tlx/tutorials/testing/test_correctness.py

# Device already set -> replace, don't stack
CUDA_VISIBLE_DEVICES=4 third_party/tlx/denoise.sh python bench.py
# becomes
CUDA_VISIBLE_DEVICES=2 third_party/tlx/denoise.sh python bench.py

Note: denoise.sh defaults to device 4 when CUDA_VISIBLE_DEVICES is unset (third_party/tlx/denoise.sh:6), so always set it explicitly when wrapping a benchmark with denoise.sh after a failure.

© facebookexperimental, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/debug-failing-gpu of facebookexperimental/triton.

Open the folder on GitHubat commit 6f3dd70

Compare with similar skills

Debug Failing GPU next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Failing GPU compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Failing GPU this skillfacebookexperimental/triton201—~709Automated safety check: PassMIT
ONNX Runtime GPU Transformers Testsmicrosoft/onnxruntime22k—~2.9kAutomated safety check: PassMIT
Jetson Video SetupNVIDIA/skills3.6k1 repos~2.4kAutomated safety check: NotesApache-2.0
Langchain CI Integrationjeremylongshore/tons-of-skills-marketplace2.8k—~4.4kAutomated safety check: PassMIT
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Running Testsbrendanhasz/probflow175—~657Automated safety check: PassMIT

Similar skills

  • Official

    Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

    22k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Jetson Video Setup

    NVIDIA/skills

    Official

    A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

    3.6k GitHub starsUsed in 1 repo~2.4k tokens
    AI & LLM EngineeringAuto-check: notes
  • Langchain CI Integration

    jeremylongshore/tons-of-skills-marketplace

    Wire LangChain 1.0 / LangGraph 1.0 tests into a GitHub Actions pipeline — unit tests with FakeListChatModel, VCR-gated integration tests, warning-filter policy, and eval-regression merge gates.

    2.8k GitHub stars~4.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Running Tests

    brendanhasz/probflow

    Run Python unit test suites strictly using the uv package manager and pytest.

    175 GitHub stars~657 tokensUpdated 13 days ago
    Testing & QAAuto-check passed
  • Add Jit Kernel

    guqiong96/Lsglang

    Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

    144 GitHub starsUsed in 1 repo~10k tokens
    AI & LLM EngineeringAuto-check passed

More from facebookexperimental/triton

All 18 skills in this repo
  • Amd Att Trace

    facebookexperimental/triton

    Official

    Collect, validate, package, and inspect rocprofv3 Advanced Thread Trace bundles for AMD GPU kernels.

    201 GitHub stars~733 tokensUpdated today
    Auto-check passed
  • Ir Override Ablation

    facebookexperimental/triton

    Official

    Design and run Triton TTGIR debugging ablations using iroverride.

    201 GitHub stars~978 tokensUpdated today
    Auto-check passed
  • Tlx Kernel Optimization Agent

    facebookexperimental/triton

    Official

    Execute the TLX Kernel Optimization Agent CLI on a Triton or TLX kernel.

    201 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Compute Sanitizer

    facebookexperimental/triton

    Official

    Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs.

    201 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Ir Debugging

    facebookexperimental/triton

    Official

    Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

    201 GitHub stars~644 tokensUpdated today
    Auto-check passed
  • Kernel Perf Testing

    facebookexperimental/triton

    Official

    Run TLX kernel performance benchmarks on Hopper, Blackwell, and AMD (gfx950/CDNA4, gfx1250) GPUs.

    201 GitHub stars~1.1k tokensUpdated today
    Auto-check passed

Questions about Debug Failing GPU

What does Debug Failing GPU do?

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. Debug Failing GPU is an agent skill from facebookexperimental/triton, published by the product's own GitHub organization. Recover from GPU-busy / GPU-unavailable failures.

When should I use Debug Failing GPU?

Debug Failing GPU fits situations like: A command (pytest; A TLX/Triton kernel run; A benchmark) fails with errors indicating the GPU is busy; unavailable — e.g.

How do I install Debug Failing GPU in Claude Code?

Run `npx skills add facebookexperimental/triton --skill debug-failing-gpu -a claude-code`. Or copy the skill folder (.claude/skills/debug-failing-gpu in facebookexperimental/triton) into .claude/skills/debug-failing-gpu in your project. Claude Code loads it when a task matches its description.

How do I install Debug Failing GPU in Codex?

Run `npx skills add facebookexperimental/triton --skill debug-failing-gpu -a codex`. Or copy the skill folder (.claude/skills/debug-failing-gpu in facebookexperimental/triton) into .agents/skills/debug-failing-gpu in your project. Codex loads it when a task matches its description.

Can I use Debug Failing GPU in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add facebookexperimental/triton --skill debug-failing-gpu -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-failing-gpu, .gemini/skills/debug-failing-gpu, .github/skills/debug-failing-gpu and .opencode/skills/debug-failing-gpu in your project.

What does Debug Failing GPU need to run?

Going by SKILL.md and its folder, Debug Failing GPU needs the command-line tools its instructions call (bash and pytest). Our summary lists: Python 3.

Does Debug Failing GPU access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Debug Failing GPU safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debug Failing GPU use?

Debug Failing GPU is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Failing GPU use?

About 709 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Failing GPU?

Skills that share tags, products or a category with Debug Failing GPU: ONNX Runtime GPU Transformers Tests (microsoft/onnxruntime, 22k stars), Jetson Video Setup (NVIDIA/skills, 3.6k stars), Langchain CI Integration (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Failing GPU?

facebookexperimental (a GitHub organization, an official publisher) maintains it in facebookexperimental/triton, which has 201 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on October 10, 2026.

Source: facebookexperimental/triton on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.