Make Op Verify
CVCUDA/CV-CUDA
Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).
Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.
$ npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install slowlyC/agent-gpu-skills cuda-skill --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cuda-skill .claude/skills/cuda-skill && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cuda-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/cuda-skill into .claude/skills/cuda-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-skill", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/cuda-skillType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install slowlyC/agent-gpu-skills cuda-skill --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/cuda-skill .agents/skills/cuda-skill && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cuda-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/cuda-skill into .agents/skills/cuda-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-skill", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install slowlyC/agent-gpu-skills cuda-skill --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/cuda-skill .cursor/skills/cuda-skill && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cuda-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/cuda-skill into .cursor/skills/cuda-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-skill", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/slowlyC/agent-gpu-skills.git --path skills/cuda-skill--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install slowlyC/agent-gpu-skills cuda-skill --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/cuda-skill .gemini/skills/cuda-skill && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cuda-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/cuda-skill into .gemini/skills/cuda-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-skill", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install slowlyC/agent-gpu-skills cuda-skillInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/cuda-skill .github/skills/cuda-skill && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cuda-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/cuda-skill into .github/skills/cuda-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-skill", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install slowlyC/agent-gpu-skills cuda-skill --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/cuda-skill .opencode/skills/cuda-skill && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cuda-skill" agent skill from https://github.com/slowlyC/agent-gpu-skills/tree/main/skills/cuda-skill into .opencode/skills/cuda-skill/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-skill", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cuda-skillQuery current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.
Cuda Skill is an agent skill from slowlyC/agent-gpu-skills. Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Use for direct CUDA C++ or PTX work, and for framework tasks only when they need NVIDIA ISA, API, architecture, or tool facts. Triggers include inline PTX, WMMA, WGMMA, TMA, tcgen05, mbarrier, fabric operations, CUDA APIs and Graphs, memory ordering, compute capability, Ampere, Hopper, Blackwell, Rubin, nsys, ncu, and compute-sanitizer.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 900 other files, including reference files (for example `references/MANIFEST.md`, `references/best-practices-guide/1-overview.md` and `references/best-practices-guide/10-memory-optimizations.md`).
It sits in Development. It works with CUDA, NVIDIA AI Platform and C++. The licence is MIT.
Read from SKILL.md and the folder at commit ae02d07. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cuda Skill loads about 1.8k tokens when it runs, and up to ~2M if it reads all its reference files. Until then it costs about 123 tokens; SKILL.md has 812 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from slowlyC/agent-gpu-skills at commit ae02d07, republished under its MIT licence (© slowlyC). 812 words, ~1,824 tokens.
.claude/skills/cuda-skill/SKILL.md (or your agent's skills folder). This skill also uses 898 other files; get the full folder from GitHub.Use this skill as the source of truth for CUDA, PTX, NVIDIA GPU architecture, and NVIDIA profiling or debugging tools. Prefer the local official-document snapshots, then verify against NVIDIA's current online documentation when a fact is version-sensitive or absent locally.
For DSL- or library-specific implementation, use the corresponding skill first:
Add this skill when those tasks require CUDA API, PTX ISA, architecture, or NVIDIA tool facts.
Resolve the directory containing this SKILL.md, then use its references/ child. Do not assume a Cursor, Claude, or Codex-specific install path.
In examples below, set a task-scoped variable to the resolved absolute path:
CUDA_REFS=/absolute/path/to/cuda-skill/referencesRead MANIFEST.md before making version claims. It records the snapshot version, source URL, and document inventory.
| Question | Primary source |
|---|---|
| PTX syntax, semantics, ISA or target requirements | ptx-docs/ |
| CUDA Runtime functions, errors, and structs | cuda-runtime-docs/ |
| CUDA Driver functions, contexts, modules, VMM | cuda-driver-docs/ |
| CUDA programming model and feature behavior | cuda-guide/ |
| General CUDA optimization guidance | best-practices-guide/ |
| Nsight Compute metrics, sections, and CLI | ncu-docs/, ncu-guide.md |
| Nsight Systems tracing and CLI | nsys-docs/, nsys-guide.md |
| Correctness tools and cuda-gdb | debugging-tools.md |
| NVTX instrumentation | nvtx-patterns.md |
| Frequent performance mistakes | performance-traps.md |
The short guide files are search maps, not substitutes for the full official snapshots.
Start with file discovery. Do not load a large chapter or the whole specification when a focused page exists.
# Discover focused PTX pages.
rg -l -i 'wgmma\.mma_async' "$CUDA_REFS/ptx-docs"
# Read the relevant lines with context.
rg -n -C 12 'Target ISA Notes|PTX ISA Notes|wgmma\.mma_async' \
"$CUDA_REFS/ptx-docs/9-instruction-set"
# Runtime and Driver API lookup.
rg -l 'cudaStreamSynchronize' "$CUDA_REFS/cuda-runtime-docs"
rg -l 'cuMemMap' "$CUDA_REFS/cuda-driver-docs"
# Programming and optimization concepts.
rg -l -i 'thread block cluster' "$CUDA_REFS/cuda-guide"
rg -l -i 'coalesc' "$CUDA_REFS/best-practices-guide"For PTX instructions, inspect all of the following before answering:
Keep these four layers separate:
PTX ISA version
→ virtual target accepted by the assembler
→ toolkit/compiler support
→ physical GPU capabilityA documented target does not by itself prove that the local toolkit accepts it or that the current machine implements it. For unreleased or preview architectures such as Rubin, verify the current official online documentation.
Search by exact symbol first, then read the containing module and related type pages.
rg -n -C 20 'cudaErrorInvalidValue' "$CUDA_REFS/cuda-runtime-docs"
rg -n -C 25 'cudaLaunchKernelEx' "$CUDA_REFS/cuda-runtime-docs"
rg -n -C 25 'cuCtxCreate' "$CUDA_REFS/cuda-driver-docs"
rg -n -C 25 'cuMemCreate|cuMemMap' "$CUDA_REFS/cuda-driver-docs"Check parameter lifetime, synchronization behavior, error propagation, version notes, and deprecation status. Do not infer Runtime API behavior from a similarly named Driver API function.
Minimize the reproducer, preserve the failing launch configuration, then use the narrowest correctness tool:
compute-sanitizer --tool memcheck ./program
compute-sanitizer --tool racecheck ./program
compute-sanitizer --tool initcheck ./program
compute-sanitizer --tool synccheck ./programUse debugging-tools.md for tool options and limitations. After a fix, rerun the original workload because sanitizer execution changes scheduling and timing.
Use Nsight Systems to locate time and overlap problems, then Nsight Compute to explain one selected kernel.
nsys profile -o report ./program
nsys stats report.nsys-rep --report cuda_gpu_kern_sum
ncu --list-sets
ncu --list-sections
ncu --query-metrics
ncu --kernel-name regex:myKernel --launch-count 1 -o report ./programMetric names, section identifiers, predefined sets, and report formats can change between releases and architectures. Discover what the active tool supports, then confirm semantics in the latest local Nsight documentation. Do not bind guidance to the machine's installed NCU version.
Base conclusions on measured evidence:
Change one hypothesis at a time and remeasure against the same baseline.
For Ampere, Hopper, Blackwell, or Rubin questions, distinguish public architecture disclosures from ISA availability. Check:
Do not identify a GPU architecture solely from a failed CUDA runtime query or a product label. Use explicit compute-capability or compilation-target evidence when available.
Always scrape into a fresh staging root. --force overwrites matching files but does not delete the output directory or unrelated files.
cd /path/to/agent-gpu-skills
uv run scripts/scrape_docs.py all \
--output-dir /tmp/cuda-docs-staging \
--force
diff -qr skills/cuda-skill/references/ptx-docs \
/tmp/cuda-docs-staging/ptx-docsReview version changes, page-count changes, renamed files, and representative instruction/API pages before merging. Do not remove obsolete live files without explicit user approval.
Run the repository validator after any update:
python3 scripts/validate_cuda_skill.pyState which document version supports the answer. Cite the focused local file and section when possible. If online verification was required, link the official NVIDIA page and label any inference. Avoid hardcoded performance thresholds unless they come from the user's measurements or a cited document.
© slowlyC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 898 other files (references) in skills/cuda-skill of slowlyC/agent-gpu-skills.
Open the folder on GitHubat commit ae02d07
Cuda Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cuda Skill this skillslowlyC/agent-gpu-skills | 169 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Make Op VerifyCVCUDA/CV-CUDA | 2.7k | — | ~433 | Automated safety check: Pass | Custom licence | |
| Review Op SupportCVCUDA/CV-CUDA | 2.7k | — | ~248 | Automated safety check: Pass | Custom licence | |
| Review Op Test CoverageCVCUDA/CV-CUDA | 2.7k | — | ~264 | Automated safety check: Pass | Custom licence | |
| Cudaq GuideNVIDIA/skills | 3.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Cuopt DeveloperNVIDIA/skills | 3.5k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 |
CVCUDA/CV-CUDA
Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).
CVCUDA/CV-CUDA
Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix.
CVCUDA/CV-CUDA
Review a CV-CUDA operator's test coverage, including C++ correctness, required cross-layout parity, correctness rigor, and the Python API surface.
NVIDIA/skills
A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance.
NVIDIA/skills
Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI).
NVIDIA/skills
Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
slowlyC/agent-gpu-skills
Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.
slowlyC/agent-gpu-skills
Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.
Works with
Categories
Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. Cuda Skill is an agent skill from slowlyC/agent-gpu-skills. Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.
Cuda Skill fits situations like: direct CUDA C++; for framework tasks only when they need NVIDIA ISA; include inline PTX; fabric operations.
Run `npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a claude-code`. Or copy the skill folder (skills/cuda-skill in slowlyC/agent-gpu-skills) into .claude/skills/cuda-skill in your project. Claude Code loads it when a task matches its description.
Run `npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a codex`. Or copy the skill folder (skills/cuda-skill in slowlyC/agent-gpu-skills) into .agents/skills/cuda-skill in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add slowlyC/agent-gpu-skills --skill cuda-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-skill, .gemini/skills/cuda-skill, .github/skills/cuda-skill and .opencode/skills/cuda-skill in your project.
SKILL.md names no scripts, command-line tools or credentials: Cuda Skill is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cuda Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2M tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Cuda Skill: Make Op Verify (CVCUDA/CV-CUDA, 2.7k stars), Review Op Support (CVCUDA/CV-CUDA, 2.7k stars), Review Op Test Coverage (CVCUDA/CV-CUDA, 2.7k stars) and Cudaq Guide (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
slowlyC (a GitHub user) maintains it in slowlyC/agent-gpu-skills, which has 169 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 8, 2026.
Source: slowlyC/agent-gpu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.