Agent skill

Neuron Nki Debugging

by uw-syfi in uw-syfi/vibesys

This skill guides debugging NKI compilation errors on Neuron hardware.

MITAuto-check passedDevelopment

Install Neuron Nki Debugging

skills CLI
$ npx skills add uw-syfi/vibesys --skill neuron-nki-debugging -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install uw-syfi/vibesys neuron-nki-debugging --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/uw-syfi/vibesys.git skills-src && mkdir -p .claude/skills && cp -r skills-src/resources/skills/neuron-agentic-development/skills/neuron-nki-debugging .claude/skills/neuron-nki-debugging && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
neuron-nki-debugging
GitHub stars
103
Token cost
~2.7k tokens
SKILL.md length
624 words
Files
8 (incl. references)
Skills in repo
15
Repo updated
First seen
Licence
MIT

At a glance

This skill guides debugging NKI compilation errors on Neuron hardware.

  • Works in 5 steps: Set Environment Variables → Apply Matching Decorator → Create Test Script → …
  • Encountering compiler error on device
  • SKILL.md covers Quick Start, Prerequisites, Platform Detection and Standard Debugging Workflow, plus 5 more sections
  • Calls python

What it does

Neuron Nki Debugging is an agent skill from uw-syfi/vibesys. This skill guides debugging NKI compilation errors on Neuron hardware. Use when encountering "compiler error on device", "debug NKI kernel", "test kernel on trn2/trn3", "neuronx-cc compilation failed", "validate kernel on hardware", "run kernel on trainium", or asking "how to debug NKI compilation errors on device".

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `references/compiler-artifacts.md`, `references/compiler-error-codes.md` and `references/compiler-flags.md`).

It sits in Development, covering Debugging. The repository describes itself as: Can AI Agents Build Bespoke Systems? The licence is MIT.

When your agent uses it

  • Encountering compiler error on device
  • Debug NKI kernel
  • Test kernel on trn2/trn3
  • Neuronx-cc compilation failed

Example prompts

  • “compiler error on device”
  • “debug NKI kernel”
  • “test kernel on trn2/trn3”
  • “/neuron-nki-debugging”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Set Environment Variables
  2. Apply Matching Decorator
  3. Create Test Script
  4. Run and Observe
  5. Validate Numerically

What it can do on your machine

Read from SKILL.md and the folder at commit c7784eb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Neuron Nki Debugging loads about 2.7k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 85 tokens; SKILL.md has 624 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~85
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from uw-syfi/vibesys at commit c7784eb, republished under its MIT licence (© uw-syfi). 624 words, ~2,692 tokens.

Download SKILL.mdSave it as .claude/skills/neuron-nki-debugging/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
neuron-nki-debugging
description
This skill guides debugging NKI compilation errors on Neuron hardware. Use when encountering "compiler error on device", "debug NKI kernel", "test kernel on trn2/trn3", "neuronx-cc compilation failed", "validate kernel on hardware", "run kernel on trainium", or asking "how to debug NKI compilation errors on device".
argument-hint
[kernel file]

Debugging NKI on Neuron Hardware

This skill provides a workflow for debugging NKI kernel compilation and execution on Trainium/Inferentia hardware.

Quick Start

Minimal working example to test kernel compilation on device:

python
import os
import torch
from torch_xla.core import xla_model as xm
import nki
import nki.language as nl
import nki.isa as nisa

os.environ["NEURON_CC_FLAGS"] = "--target trn2 --lnc 1"
os.environ["NEURON_PLATFORM_TARGET_OVERRIDE"] = "trn2"

@nki.jit
def add_kernel(a_input, b_input):
    """Element-wise addition kernel."""
    a_tile = nl.ndarray(a_input.shape, dtype=a_input.dtype, buffer=nl.sbuf)
    nisa.dma_copy(dst=a_tile, src=a_input[0:a_input.shape[0], 0:a_input.shape[1]])

    b_tile = nl.ndarray(b_input.shape, dtype=b_input.dtype, buffer=nl.sbuf)
    nisa.dma_copy(dst=b_tile, src=b_input[0:b_input.shape[0], 0:b_input.shape[1]])

    c_tile = nl.ndarray(a_input.shape, dtype=a_input.dtype, buffer=nl.sbuf)
    nisa.tensor_tensor(dst=c_tile, data1=a_tile, data2=b_tile, op=nl.add)

    c_output = nl.ndarray(a_input.shape, dtype=a_input.dtype, buffer=nl.shared_hbm)
    nisa.dma_copy(dst=c_output, src=c_tile)
    return c_output

# Get XLA device and run
device = xm.xla_device()
a = torch.ones((4, 3), dtype=torch.float16).to(device=device)
b = torch.ones((4, 3), dtype=torch.float16).to(device=device)

c = add_kernel(a, b)
print(c)  # Forces XLA compilation and execution

Prerequisites

Before running kernels on device, resolve the NKI virtual environment path:

  1. Check environment: echo $NKI_VENV_PATH
  2. If empty, read .claude/nki-dev-suite.local.md and extract nki_venv_path from YAML frontmatter
  3. If still not found, report: "NKI_VENV_PATH not configured. Set the environment variable or create .claude/nki-dev-suite.local.md with nki_venv_path in frontmatter."

Activate before running any device tests:

bash
source $NKI_VENV_PATH/bin/activate

Platform Detection

Before compilation, detect the current hardware platform:

Current platform: !neuron-ls | head -3

Platform Target Mapping
HardwareInstanceTarget FlagGeneration
Trainium 1trn1--target trn1gen2
Trainium 1ntrn1n--target trn1ngen2
Inferentia 2inf2--target inf2gen2
Trainium 2trn2--target trn2gen3
Trainium 3trn3--target trn3gen4

Match the --target flag and platform_target decorator argument to your detected hardware.

Standard Debugging Workflow

Step 1: Set Environment Variables
python
import os

# Standard debugging flags (minimal, fast compilation)
os.environ["NEURON_CC_FLAGS"] = "--target trn2 --lnc 1"
os.environ["NEURON_PLATFORM_TARGET_OVERRIDE"] = "trn2"

# Pin to a specific neuron core to avoid conflicts with concurrent sessions
os.environ["NEURON_RT_VISIBLE_CORES"] = "0"
FlagPurpose
--targetHardware platform (trn1, trn2, trn3, inf2)
--lnc 1Single NeuronCore (simplifies debugging)
NEURON_RT_VISIBLE_CORESPin to specific core(s) — prevents contention when multiple agents run concurrently

See references/compiler-flags.md for complete flag reference.

Step 2: Apply Matching Decorator
python
@nki.jit  # Must match --target and NEURON_PLATFORM_TARGET_OVERRIDE
def my_kernel(input_tensor):
    ...

The platform_target environment variable MUST match the --target in NEURON_CC_FLAGS.

Step 3: Create Test Script
python
import os
import torch
from torch_xla.core import xla_model as xm
import nki

os.environ["NEURON_CC_FLAGS"] = "--target trn2 --lnc 1"
os.environ["NEURON_PLATFORM_TARGET_OVERRIDE"] = "trn2"

@nki.jit
def kernel(input_tensor):
    # Your kernel implementation
    ...
    return output_tensor

# XLA device execution pattern
device = xm.xla_device()
input_data = torch.randn((128, 512), dtype=torch.float32).to(device=device)

output = kernel(input_data)
print(output)  # Forces XLA compilation - triggers actual compilation
Step 4: Run and Observe
bash
source $NKI_VENV_PATH/bin/activate
python your_test_script.py

Compilation errors appear in the console output. The print() statement forces XLA compilation, which triggers the neuronx-cc compiler.

Step 5: Validate Numerically

Compare device output against a CPU-computed reference using multiple complementary checks — no single metric catches all issues:

  • atol / rtol (torch.allclose): Per-element pass/fail gate
  • Maximum absolute difference: Worst-case outlier check
  • Norm of the difference tensor: Detects widespread small drift
  • Cosine similarity: Catches directional errors in high-dimensional outputs

Important: Compute references on CPU, not on the XLA device. Every XLA graph compiled on-device generates a separate NEFF file. Running reference operations (e.g., torch.matmul, torch.softmax) on the XLA device creates extra NEFFs, making it hard to identify which NEFF belongs to the NKI kernel during profiling.

python
# CORRECT: Reference computed on CPU — only the NKI kernel generates a NEFF
cpu_input = input_data.cpu()
reference_output = reference_implementation(cpu_input)
device_output = output.cpu()

# Use dtype-appropriate tolerances
assert torch.allclose(device_output, reference_output, rtol=1e-5, atol=1e-8)
python
# WRONG: Reference computed on device — generates an extra NEFF
reference_output = reference_implementation(input_data)  # Compiles to separate NEFF!

For complex kernels where the final output is wrong: decompose the kernel into logical stages and examine intermediate tensors at each boundary. Store intermediates to HBM temporarily, compare each against the matching reference stage, and binary-search for the stage that introduces the error. Once the failing stage is identified, test it with minimal input shapes (e.g., a single 128x128 tile) to isolate whether the issue is in the core logic or in tiling/boundary handling. Remove the debug stores once the issue is resolved.

Show full SKILL.md (211 more words)Show less

Compiler Artifacts Mode

For advanced debugging that preserves compiler outputs for inspection, use when you need to understand detailed compilation behavior.

When to use: "compiler artifacts", "compiler flags", "inspect compiler log"

See references/compiler-artifacts.md for:

  • Compiler debug flag configuration (--verbose, --target, --lnc)
  • Finding the compiler temp folder
  • Understanding generated artifacts (*.neff, log-neuron-cc.txt)

Error Resolution

Error Categories
Error PatternCategoryReference
NCC_EVRF*Verification errorSee references/ncc-verification-errors.md
NCC_EOOM*Out of memorySee references/ncc-memory-resource-errors.md
NCC_E* (other)Type/operation errorSee references/ncc-type-operation-errors.md
Quick Reference

See references/compiler-error-codes.md for the complete index of all 28 NCC_* error codes.

Common Error Quick Fixes
Error CodeCategoryQuick Fix
NCC_EVRF001Unsupported operatorUse alternative operator from neuronx-cc list-operators
NCC_EOOM001Memory exceededReduce batch size, use tensor/pipeline parallelism
NCC_EVRF007Instruction limitApply model parallelism
NCC_EVRF005Unsupported FP8 typeConvert to float16/bfloat16 or use gen3+ hardware
NCC_EARG001LNC configurationUse supported LNC count for target hardware
NCC_EVRF024Output tensor > 4GBReduce tensor size or use tensor parallelism

Profiling (Optional)

To capture execution traces for profiling:

python
# Add before running kernel
os.environ['NEURON_RT_INSPECT_ENABLE'] = '1'
os.environ['NEURON_RT_INSPECT_DEVICE_PROFILE'] = '1'
os.environ['NEURON_RT_INSPECT_OUTPUT_DIR'] = './output'

This captures NEFF (compiled binary) and NTFF (execution trace) files in the output directory.

Complete Example

python
import os
import torch
from torch_xla.core import xla_model as xm
import nki
import nki.language as nl
import nki.isa as nisa

# Standard debugging configuration
os.environ["NEURON_CC_FLAGS"] = "--target trn2 --lnc 1"

# Optional: Enable profiling
os.environ['NEURON_RT_INSPECT_ENABLE'] = '1'
os.environ['NEURON_RT_INSPECT_DEVICE_PROFILE'] = '1'
os.environ['NEURON_RT_INSPECT_OUTPUT_DIR'] = './output'
os.environ["NEURON_PLATFORM_TARGET_OVERRIDE"] = "trn2"

@nki.jit
def softmax_kernel(input_tensor):
    """Simple softmax along last dimension."""
    # Load input tile
    tile = nl.ndarray(input_tensor.shape, dtype=input_tensor.dtype, buffer=nl.sbuf)
    nisa.dma_copy(dst=tile, src=input_tensor)

    # Compute softmax
    exp_tile = nl.ndarray(input_tensor.shape, dtype=input_tensor.dtype, buffer=nl.sbuf)
    nisa.activation(dst=exp_tile, data=tile, op=nl.exp)

    sum_tile = nl.ndarray((input_tensor.shape[0], 1), dtype=input_tensor.dtype, buffer=nl.sbuf)
    nisa.tensor_reduce(dst=sum_tile, data=exp_tile, op=nl.add, axis=(1,))

    recip_sum = nl.ndarray((input_tensor.shape[0], 1), dtype=input_tensor.dtype, buffer=nl.sbuf)
    nisa.reciprocal(dst=recip_sum, data=sum_tile)

    result = nl.ndarray(input_tensor.shape, dtype=input_tensor.dtype, buffer=nl.sbuf)
    nisa.tensor_scalar(dst=result, data=exp_tile, op0=nl.multiply, operand0=recip_sum)

    # Store output
    output = nl.ndarray(input_tensor.shape, dtype=input_tensor.dtype, buffer=nl.shared_hbm)
    nisa.dma_copy(dst=output, src=result)
    return output

# Test execution
device = xm.xla_device()
x = torch.randn((64, 128), dtype=torch.float32).to(device=device)

y = softmax_kernel(x)
print(y)  # Triggers compilation

# Validate against PyTorch reference
reference = torch.softmax(x.cpu(), dim=-1)
assert torch.allclose(y.cpu(), reference, rtol=1e-4, atol=1e-6)
print("Validation passed!")

Configuration

Required settings:

SettingSourceDescription
nki_venv_path.claude/nki-dev-suite.local.md or NKI_VENV_PATHPython venv with neuronx packages

Related skills:

SkillUse When
/neuron-nki-profilingProfile kernel performance
/neuron-nki-docsLook up API documentation and error codes

© uw-syfi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in resources/skills/neuron-agentic-development/skills/neuron-nki-debugging of uw-syfi/vibesys.

  • SKILL.md
  • references/compiler-artifacts.md
  • references/compiler-error-codes.md
  • references/compiler-flags.md
  • references/ncc-memory-resource-errors.md
  • references/ncc-type-operation-errors.md
  • references/ncc-verification-errors.md
  • references/neuron-core-isolation.md

Open the folder on GitHubat commit c7784eb

Compare with similar skills

Neuron Nki Debugging next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Neuron Nki Debugging compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Neuron Nki Debugging this skilluw-syfi/vibesys103—~2.7kAutomated safety check: PassMIT
Trellis Session Insightmindfold-ai/Trellis15k4 repos~1.7kAutomated safety check: PassAGPL-3.0
Native Data FetchingCherryHQ/cherry-studio-app4k6 repos~2.9kAutomated safety check: NotesMIT
Debugging Executionsn8n-io/n8n207k—~2.6kAutomated safety check: PassCustom licence
Aoti Debugpytorch/pytorch104k1 repos~1.7kAutomated safety check: PassCustom licence
Herdr Throwaway Reproductionherdrdev/herdr43k—~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Trellis Session Insight

    mindfold-ai/Trellis

    Reach into past AI conversation history through the trellis mem CLI.

    15k GitHub starsUsed in 4 repos~1.7k tokens
    DevelopmentAuto-check passed
  • Native Data Fetching

    CherryHQ/cherry-studio-app

    A skill your agent uses when implementing or debugging ANY network request, API call, or data fetching.

    4k GitHub starsUsed in 6 repos~2.9k tokens
    DevelopmentAuto-check: notes
  • Official

    Debug failed or wrong-output workflow executions using executions tools.

    207k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • Runs a disposable, uniquely named Herdr session inside an existing one so runtime, pane, terminal or API bugs can be reproduced without touching the main session.

    43k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Systematic Debugging

    ultralisp/ultralisp

    A skill your agent uses when encountering any bug, test failure, or unexpected behavior, before proposing fixes

    258 GitHub starsUsed in 51 repos~2.4k tokens
    DevelopmentAuto-check passed

More from uw-syfi/vibesys

All 15 skills in this repo
  • Neuron Nki Profiling

    uw-syfi/vibesys

    This skill guides using the cli to generate NKI kernel profiles (NEFF + NTFF pairs) to analyze performance on Neuron hardware.

    103 GitHub stars~3.2k tokensUpdated today
    Auto-check passed
  • Neuron Nki Docs

    uw-syfi/vibesys

    Research NKI documentation for API lookups, tutorials, error codes, architecture, and optimization guides.

    103 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Query and analyze NKI kernel profile data from neuron-explorer parquet files.

    103 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Neuron Nki Writing

    uw-syfi/vibesys

    Guide for writing and modifying NKI kernels. An agent skill from uw-syfi/vibesys.

    103 GitHub stars~5k tokensUpdated today
    Auto-check passed
  • Triage PRs

    uw-syfi/vibesys

    Triage the open pull requests of the VibeSys repository. An agent skill from uw-syfi/vibesys.

    103 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Open PR

    uw-syfi/vibesys

    Prepare and open VibeSys pull requests from local repo changes.

    103 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Neuron Nki Debugging

What does Neuron Nki Debugging do?

This skill guides debugging NKI compilation errors on Neuron hardware. Neuron Nki Debugging is an agent skill from uw-syfi/vibesys. This skill guides debugging NKI compilation errors on Neuron hardware.

When should I use Neuron Nki Debugging?

Neuron Nki Debugging fits situations like: encountering compiler error on device; debug NKI kernel; test kernel on trn2/trn3; neuronx-cc compilation failed.

How do I install Neuron Nki Debugging in Claude Code?

Run `npx skills add uw-syfi/vibesys --skill neuron-nki-debugging -a claude-code`. Or copy the skill folder (resources/skills/neuron-agentic-development/skills/neuron-nki-debugging in uw-syfi/vibesys) into .claude/skills/neuron-nki-debugging in your project. Claude Code loads it when a task matches its description.

How do I install Neuron Nki Debugging in Codex?

Run `npx skills add uw-syfi/vibesys --skill neuron-nki-debugging -a codex`. Or copy the skill folder (resources/skills/neuron-agentic-development/skills/neuron-nki-debugging in uw-syfi/vibesys) into .agents/skills/neuron-nki-debugging in your project. Codex loads it when a task matches its description.

Can I use Neuron Nki Debugging in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add uw-syfi/vibesys --skill neuron-nki-debugging -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/neuron-nki-debugging, .gemini/skills/neuron-nki-debugging, .github/skills/neuron-nki-debugging and .opencode/skills/neuron-nki-debugging in your project.

What does Neuron Nki Debugging need to run?

Going by SKILL.md and its folder, Neuron Nki Debugging needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Neuron Nki Debugging access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Neuron Nki Debugging safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Neuron Nki Debugging use?

Neuron Nki Debugging is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Neuron Nki Debugging use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Neuron Nki Debugging?

Skills that share tags, products or a category with Neuron Nki Debugging: Trellis Session Insight (mindfold-ai/Trellis, 15k stars), Native Data Fetching (CherryHQ/cherry-studio-app, 4k stars), Debugging Executions (n8n-io/n8n, 207k stars) and Aoti Debug (pytorch/pytorch, 104k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Neuron Nki Debugging?

uw-syfi (a GitHub organization) maintains it in uw-syfi/vibesys, which has 103 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 9, 2026.

Source: uw-syfi/vibesys on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.