CUDA kernel development, debugging, and performance optimization for Claude Code.

No licenceAuto-check passedDevelopment

Install Cuda

skills CLI
$ npx skills add technillogue/ptx-isa-markdown --skill cuda -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install technillogue/ptx-isa-markdown cuda --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/technillogue/ptx-isa-markdown.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cuda_skill .claude/skills/cuda && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cuda
GitHub stars
229
Token cost
~2.5k tokens
SKILL.md length
828 words
Files
665 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
None found

At a glance

CUDA kernel development, debugging, and performance optimization for Claude Code.

  • Works in 5 steps: Reproduce minimally — Isolate the… → Add printf — Before any tool, add printf… → Run compute-sanitizer — Catch memory… → …
  • Optimizing CUDA code
  • SKILL.md covers Core Philosophy, Debugging Workflow, Performance Optimization… and Compilation Reference, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cuda is an agent skill from technillogue/ptx-isa-markdown. CUDA kernel development, debugging, and performance optimization for Claude Code. Use when writing, debugging, or optimizing CUDA code, GPU kernels, or parallel algorithms. Covers non-interactive profiling with nsys/ncu, debugging with cuda-gdb/compute-sanitizer, binary inspection with cuobjdump, and performance analysis workflows. Triggers on CUDA, GPU programming, kernel optimization, nsys, ncu, cuda-gdb, compute-sanitizer, PTX, GPU profiling, parallel performance.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 667 other files, including reference files (for example `references/cuda-driver-docs/INDEX.md`, `references/cuda-driver-docs/data-structures/structcu__dev__sm__resource__group__params.md` and `references/cuda-driver-docs/data-structures/structcuaccesspolicywindow__v1.md`).

It sits in Development, covering Performance optimization, GPU and accelerator computing and Debugging. It works with CUDA. The repository describes itself as: PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.

When your agent uses it

  • Optimizing CUDA code
  • Parallel algorithms
  • GPU programming
  • Kernel optimization

Example prompts

  • “/cuda”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Reproduce minimally — Isolate the failing kernel with smallest possible input
  2. Add printf — Before any tool, add printf in device code to trace execution
  3. Run compute-sanitizer — Catch memory errors non-interactively
  4. If still stuck, try cuda-gdb non-interactively for backtrace
  5. When tools fail — Minimize the diff between working and broken code. Read it. The bug is in the diff.

What it can do on your machine

Read from SKILL.md and the folder at commit 64cfba5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and cuda).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cuda loads about 2.5k tokens when it runs, and up to ~809k if it reads all its reference files. Until then it costs about 119 tokens; SKILL.md has 828 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~809k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 828 words (~2,491 tokens).

“Measure before guessing. GPU performance is deeply counterintuitive. Profile first, hypothesize second, change third, verify fourth.”

— opening of SKILL.md by technillogue
name
cuda

Read the full SKILL.md on GitHub

Files

SKILL.md and 664 other files (references) in cuda_skill of technillogue/ptx-isa-markdown.

  • SKILL.md
  • references/cuda-driver-docs/INDEX.md
  • references/cuda-driver-docs/data-structures/structcu__dev__sm__resource__group__params.md
  • references/cuda-driver-docs/data-structures/structcuaccesspolicywindow__v1.md
  • references/cuda-driver-docs/data-structures/structcuarraymapinfo__v1.md
  • references/cuda-driver-docs/data-structures/structcuasyncnotificationinfo.md
  • references/cuda-driver-docs/data-structures/structcucheckpointcheckpointargs.md
  • references/cuda-driver-docs/data-structures/structcucheckpointgpupair.md
  • references/cuda-driver-docs/data-structures/structcucheckpointlockargs.md
  • references/cuda-driver-docs/data-structures/structcucheckpointrestoreargs.md
  • references/cuda-driver-docs/data-structures/structcucheckpointunlockargs.md
  • references/cuda-driver-docs/data-structures/structcuctxcigparam.md
  • references/cuda-driver-docs/data-structures/structcuctxcreateparams.md
  • references/cuda-driver-docs/data-structures/structcuda__array3d__descriptor__v2.md
  • references/cuda-driver-docs/data-structures/structcuda__array__descriptor__v2.md
  • references/cuda-driver-docs/data-structures/structcuda__array__memory__requirements__v1.md
  • references/cuda-driver-docs/data-structures/structcuda__array__sparse__properties__v1.md
  • references/cuda-driver-docs/data-structures/structcuda__batch__mem__op__node__params__v1.md
  • … and 647 more

Open the folder on GitHubat commit 64cfba5

Compare with similar skills

Cuda next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cuda compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cuda this skilltechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
Debug Distributed Hangsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0
CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime22k—~1.3kAutomated safety check: PassMIT
The Art of Debuggingstas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.0
Torch Profiler Layer TrackBBuf/AI-Infra-Auto-Driven-SKILLS925—~2kAutomated safety check: PassNone
Cppcrazyguitar/cppcheatsheet290—~1.8kAutomated safety check: PassMIT

Similar skills

  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 2 days ago
    DevelopmentAuto-check: notes
  • Torch Profiler Layer Track

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

    925 GitHub stars~2k tokensUpdated 4 days ago
    DevelopmentAuto-check passed
  • Cpp

    crazyguitar/cppcheatsheet

    Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics.

    290 GitHub stars~1.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Cuda

What does Cuda do?

CUDA kernel development, debugging, and performance optimization for Claude Code. Cuda is an agent skill from technillogue/ptx-isa-markdown. CUDA kernel development, debugging, and performance optimization for Claude Code.

When should I use Cuda?

Cuda fits situations like: optimizing CUDA code; parallel algorithms; GPU programming; kernel optimization.

How do I install Cuda in Claude Code?

Run `npx skills add technillogue/ptx-isa-markdown --skill cuda -a claude-code`. Or copy the skill folder (cuda_skill in technillogue/ptx-isa-markdown) into .claude/skills/cuda in your project. Claude Code loads it when a task matches its description.

How do I install Cuda in Codex?

Run `npx skills add technillogue/ptx-isa-markdown --skill cuda -a codex`. Or copy the skill folder (cuda_skill in technillogue/ptx-isa-markdown) into .agents/skills/cuda in your project. Codex loads it when a task matches its description.

Can I use Cuda in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add technillogue/ptx-isa-markdown --skill cuda -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda, .gemini/skills/cuda, .github/skills/cuda and .opencode/skills/cuda in your project.

What does Cuda need to run?

SKILL.md names no scripts, command-line tools or credentials: Cuda is instructions for the agent only.

Does Cuda access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cuda safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cuda use?

No licence was found for Cuda or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Cuda use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 806k tokens, read only when the agent opens those files.

What are the alternatives to Cuda?

Skills that share tags, products or a category with Cuda: Debug Distributed Hang (sgl-project/sglang, 37k stars), CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars), The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars) and Torch Profiler Layer Track (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cuda?

technillogue (a GitHub user) maintains it in technillogue/ptx-isa-markdown, which has 229 GitHub stars. The repository was last updated on December 24, 2025.

Source: technillogue/ptx-isa-markdown on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.