Agent skill

Cuda Skill

by slowlyC in slowlyC/agent-gpu-skills

Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.

MITAuto-check passedDevelopment

Install Cuda Skill

skills CLI
$ npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install slowlyC/agent-gpu-skills cuda-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cuda-skill .claude/skills/cuda-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cuda-skill
GitHub stars
169
Token cost
~1.8k tokens
SKILL.md length
812 words
Files
899 (incl. references)
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.

  • Direct CUDA C++
  • SKILL.md covers Locate the references, Source routing, Query workflow and CUDA API lookup, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • For framework tasks only when they need NVIDIA ISA

What it does

Cuda Skill is an agent skill from slowlyC/agent-gpu-skills. Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Use for direct CUDA C++ or PTX work, and for framework tasks only when they need NVIDIA ISA, API, architecture, or tool facts. Triggers include inline PTX, WMMA, WGMMA, TMA, tcgen05, mbarrier, fabric operations, CUDA APIs and Graphs, memory ordering, compute capability, Ampere, Hopper, Blackwell, Rubin, nsys, ncu, and compute-sanitizer.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 900 other files, including reference files (for example `references/MANIFEST.md`, `references/best-practices-guide/1-overview.md` and `references/best-practices-guide/10-memory-optimizations.md`).

It sits in Development. It works with CUDA, NVIDIA AI Platform and C++. The licence is MIT.

When your agent uses it

  • Direct CUDA C++
  • For framework tasks only when they need NVIDIA ISA
  • Include inline PTX
  • Fabric operations

Example prompts

  • “/cuda-skill”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit ae02d07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cuda Skill loads about 1.8k tokens when it runs, and up to ~2M if it reads all its reference files. Until then it costs about 123 tokens; SKILL.md has 812 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2M

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from slowlyC/agent-gpu-skills at commit ae02d07, republished under its MIT licence (© slowlyC). 812 words, ~1,824 tokens.

Download SKILL.mdSave it as .claude/skills/cuda-skill/SKILL.md (or your agent's skills folder). This skill also uses 898 other files; get the full folder from GitHub.
name
cuda-skill
description
Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Use for direct CUDA C++ or PTX work, and for framework tasks only when they need NVIDIA ISA, API, architecture, or tool facts. Triggers include inline PTX, WMMA, WGMMA, TMA, tcgen05, mbarrier, fabric operations, CUDA APIs and Graphs, memory ordering, compute capability, Ampere, Hopper, Blackwell, Rubin, nsys, ncu, and compute-sanitizer.

NVIDIA CUDA Reference

Use this skill as the source of truth for CUDA, PTX, NVIDIA GPU architecture, and NVIDIA profiling or debugging tools. Prefer the local official-document snapshots, then verify against NVIDIA's current online documentation when a fact is version-sensitive or absent locally.

For DSL- or library-specific implementation, use the corresponding skill first:

  • Triton or Gluon kernel code: triton-skill
  • CUTLASS, CuTe, or CuTeDSL code: cutlass-skill

Add this skill when those tasks require CUDA API, PTX ISA, architecture, or NVIDIA tool facts.

Locate the references

Resolve the directory containing this SKILL.md, then use its references/ child. Do not assume a Cursor, Claude, or Codex-specific install path.

In examples below, set a task-scoped variable to the resolved absolute path:

CUDA_REFS=/absolute/path/to/cuda-skill/references

Read MANIFEST.md before making version claims. It records the snapshot version, source URL, and document inventory.

Source routing

QuestionPrimary source
PTX syntax, semantics, ISA or target requirementsptx-docs/
CUDA Runtime functions, errors, and structscuda-runtime-docs/
CUDA Driver functions, contexts, modules, VMMcuda-driver-docs/
CUDA programming model and feature behaviorcuda-guide/
General CUDA optimization guidancebest-practices-guide/
Nsight Compute metrics, sections, and CLIncu-docs/, ncu-guide.md
Nsight Systems tracing and CLInsys-docs/, nsys-guide.md
Correctness tools and cuda-gdbdebugging-tools.md
NVTX instrumentationnvtx-patterns.md
Frequent performance mistakesperformance-traps.md

The short guide files are search maps, not substitutes for the full official snapshots.

Query workflow

Start with file discovery. Do not load a large chapter or the whole specification when a focused page exists.

# Discover focused PTX pages.
rg -l -i 'wgmma\.mma_async' "$CUDA_REFS/ptx-docs"

# Read the relevant lines with context.
rg -n -C 12 'Target ISA Notes|PTX ISA Notes|wgmma\.mma_async' \
  "$CUDA_REFS/ptx-docs/9-instruction-set"

# Runtime and Driver API lookup.
rg -l 'cudaStreamSynchronize' "$CUDA_REFS/cuda-runtime-docs"
rg -l 'cuMemMap' "$CUDA_REFS/cuda-driver-docs"

# Programming and optimization concepts.
rg -l -i 'thread block cluster' "$CUDA_REFS/cuda-guide"
rg -l -i 'coalesc' "$CUDA_REFS/best-practices-guide"

For PTX instructions, inspect all of the following before answering:

  • instruction syntax and operands;
  • semantic description and memory ordering;
  • PTX ISA introduction version;
  • target ISA or sm_* requirements;
  • architecture-specific restrictions and undefined behavior.

Keep these four layers separate:

PTX ISA version
  → virtual target accepted by the assembler
  → toolkit/compiler support
  → physical GPU capability

A documented target does not by itself prove that the local toolkit accepts it or that the current machine implements it. For unreleased or preview architectures such as Rubin, verify the current official online documentation.

CUDA API lookup

Search by exact symbol first, then read the containing module and related type pages.

rg -n -C 20 'cudaErrorInvalidValue' "$CUDA_REFS/cuda-runtime-docs"
rg -n -C 25 'cudaLaunchKernelEx' "$CUDA_REFS/cuda-runtime-docs"
rg -n -C 25 'cuCtxCreate' "$CUDA_REFS/cuda-driver-docs"
rg -n -C 25 'cuMemCreate|cuMemMap' "$CUDA_REFS/cuda-driver-docs"

Check parameter lifetime, synchronization behavior, error propagation, version notes, and deprecation status. Do not infer Runtime API behavior from a similarly named Driver API function.

Debugging workflow

Minimize the reproducer, preserve the failing launch configuration, then use the narrowest correctness tool:

compute-sanitizer --tool memcheck ./program
compute-sanitizer --tool racecheck ./program
compute-sanitizer --tool initcheck ./program
compute-sanitizer --tool synccheck ./program

Use debugging-tools.md for tool options and limitations. After a fix, rerun the original workload because sanitizer execution changes scheduling and timing.

Show full SKILL.md (278 more words)Show less

Profiling workflow

Use Nsight Systems to locate time and overlap problems, then Nsight Compute to explain one selected kernel.

nsys profile -o report ./program
nsys stats report.nsys-rep --report cuda_gpu_kern_sum

ncu --list-sets
ncu --list-sections
ncu --query-metrics
ncu --kernel-name regex:myKernel --launch-count 1 -o report ./program

Metric names, section identifiers, predefined sets, and report formats can change between releases and architectures. Discover what the active tool supports, then confirm semantics in the latest local Nsight documentation. Do not bind guidance to the machine's installed NCU version.

Base conclusions on measured evidence:

  • timeline placement, launch gaps, synchronization, and CPU/GPU overlap from Nsight Systems;
  • achieved throughput, instruction mix, stalls, memory traffic, occupancy, and source correlation from Nsight Compute;
  • compiler resource usage from ptxas -v or the build log.

Change one hypothesis at a time and remeasure against the same baseline.

Architecture questions

For Ampere, Hopper, Blackwell, or Rubin questions, distinguish public architecture disclosures from ISA availability. Check:

  • cuda-guide/05-appendices/compute-capabilities.md;
  • the instruction's PTX ISA and target notes;
  • ptx-docs/13-release-notes/;
  • current NVIDIA architecture or CUDA release documentation when local snapshots do not cover the claim.

Do not identify a GPU architecture solely from a failed CUDA runtime query or a product label. Use explicit compute-capability or compilation-target evidence when available.

Updating the snapshots

Always scrape into a fresh staging root. --force overwrites matching files but does not delete the output directory or unrelated files.

cd /path/to/agent-gpu-skills
uv run scripts/scrape_docs.py all \
  --output-dir /tmp/cuda-docs-staging \
  --force

diff -qr skills/cuda-skill/references/ptx-docs \
  /tmp/cuda-docs-staging/ptx-docs

Review version changes, page-count changes, renamed files, and representative instruction/API pages before merging. Do not remove obsolete live files without explicit user approval.

Run the repository validator after any update:

python3 scripts/validate_cuda_skill.py

Answer quality

State which document version supports the answer. Cite the focused local file and section when possible. If online verification was required, link the official NVIDIA page and label any inference. Avoid hardcoded performance thresholds unless they come from the user's measurements or a cited document.

© slowlyC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 898 other files (references) in skills/cuda-skill of slowlyC/agent-gpu-skills.

  • SKILL.md
  • references/MANIFEST.md
  • references/best-practices-guide/1-overview.md
  • references/best-practices-guide/10-memory-optimizations.md
  • references/best-practices-guide/10.1-data-transfer-between-host-and-device.md
  • references/best-practices-guide/10.2-device-memory-spaces.md
  • references/best-practices-guide/10.3-allocation.md
  • references/best-practices-guide/10.4-numa-best-practices.md
  • references/best-practices-guide/11-execution-configuration-optimizations.md
  • references/best-practices-guide/11.1-occupancy.md
  • references/best-practices-guide/11.2-hiding-register-dependencies.md
  • references/best-practices-guide/11.3-thread-and-block-heuristics.md
  • references/best-practices-guide/11.4-effects-of-shared-memory.md
  • references/best-practices-guide/11.5-concurrent-kernel-execution.md
  • references/best-practices-guide/11.6-multiple-contexts.md
  • references/best-practices-guide/12-instruction-optimization.md
  • references/best-practices-guide/12.1-arithmetic-instructions.md
  • references/best-practices-guide/12.2-memory-instructions.md
  • references/best-practices-guide/13-control-flow.md
  • … and 880 more

Open the folder on GitHubat commit ae02d07

Compare with similar skills

Cuda Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cuda Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cuda Skill this skillslowlyC/agent-gpu-skills169—~1.8kAutomated safety check: PassMIT
Make Op VerifyCVCUDA/CV-CUDA2.7k—~433Automated safety check: PassCustom licence
Review Op SupportCVCUDA/CV-CUDA2.7k—~248Automated safety check: PassCustom licence
Review Op Test CoverageCVCUDA/CV-CUDA2.7k—~264Automated safety check: PassCustom licence
Cudaq GuideNVIDIA/skills3.5k—~1.3kAutomated safety check: PassApache-2.0
Cuopt DeveloperNVIDIA/skills3.5k—~3.2kAutomated safety check: NotesApache-2.0

Similar skills

  • Make Op Verify

    CVCUDA/CV-CUDA

    Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).

    2.7k GitHub stars~433 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Review Op Support

    CVCUDA/CV-CUDA

    Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix.

    2.7k GitHub stars~248 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Review a CV-CUDA operator's test coverage, including C++ correctness, required cross-layout parity, correctness rigor, and the Python API surface.

    2.7k GitHub stars~264 tokensUpdated 21 days ago
    Testing & QAAuto-check passed
  • Cudaq Guide

    NVIDIA/skills

    Official

    A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance.

    3.5k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cuopt Developer

    NVIDIA/skills

    Official

    Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI).

    3.5k GitHub stars~3.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Official

    Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment.

    3.5k GitHub stars~2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from slowlyC/agent-gpu-skills

  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Tilelang Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

    169 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Triton Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed

Categories

Questions about Cuda Skill

What does Cuda Skill do?

Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Cuda Skill is an agent skill from slowlyC/agent-gpu-skills. Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.

When should I use Cuda Skill?

Cuda Skill fits situations like: direct CUDA C++; for framework tasks only when they need NVIDIA ISA; include inline PTX; fabric operations.

How do I install Cuda Skill in Claude Code?

Run `npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a claude-code`. Or copy the skill folder (skills/cuda-skill in slowlyC/agent-gpu-skills) into .claude/skills/cuda-skill in your project. Claude Code loads it when a task matches its description.

How do I install Cuda Skill in Codex?

Run `npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a codex`. Or copy the skill folder (skills/cuda-skill in slowlyC/agent-gpu-skills) into .agents/skills/cuda-skill in your project. Codex loads it when a task matches its description.

Can I use Cuda Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-skill, .gemini/skills/cuda-skill, .github/skills/cuda-skill and .opencode/skills/cuda-skill in your project.

What does Cuda Skill need to run?

SKILL.md names no scripts, command-line tools or credentials: Cuda Skill is instructions for the agent only. Our summary lists: Python 3.

Does Cuda Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cuda Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cuda Skill use?

Cuda Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cuda Skill use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2M tokens, read only when the agent opens those files.

What are the alternatives to Cuda Skill?

Skills that share tags, products or a category with Cuda Skill: Make Op Verify (CVCUDA/CV-CUDA, 2.7k stars), Review Op Support (CVCUDA/CV-CUDA, 2.7k stars), Review Op Test Coverage (CVCUDA/CV-CUDA, 2.7k stars) and Cudaq Guide (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cuda Skill?

slowlyC (a GitHub user) maintains it in slowlyC/agent-gpu-skills, which has 169 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 8, 2026.

Source: slowlyC/agent-gpu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.