Language
CUDA agent skills for Claude Code, Codex and other agents.
- skills
- 280
- official
- 90
- Type
- Language
- Website
- developer.nvidia.com
- Official GitHub
- NVIDIA
CUDA skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks. | NVIDIA/ | 3.4k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 3 | Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 4 | Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component. | microsoft/ | 22k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 5 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 6 | Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. | huggingface/ | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 7 | Diagnose and fix Cosmos3 environment, installation, and runtime errors. | NVIDIA/ | 556 | — | ~1.3k | Automated safety check: Notes | Unknown | 11 days ago |
| 8 | Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK). | intel/ | 1.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill. | NVIDIA/ | 137 | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 14 days ago |
| 10 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is… | NVIDIA/ | 137 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 14 days ago |
| 12 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 13 | Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired… | NVIDIA/ | 556 | — | ~2.7k | Automated safety check: Pass | Unknown | 11 days ago |
| 14 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 15 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 16 | A skill your agent uses when composing the Search(...) call and calling .start(). | NVIDIA/ | 137 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | 14 days ago |
| 17 | Guide users through Cosmos3 installation, environment setup, checkpoint downloading, and verification. | NVIDIA/ | 556 | — | ~1.1k | Automated safety check: Notes | Unknown | 11 days ago |
| 18 | Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. | facebookexperimental/ | 201 | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 19 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 912 | 1 repo | ~910 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 21 | Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests. | intel/ | 115 | — | ~917 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 22 | Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining. | NVIDIA/ | 3.5k | — | ~2k | Automated safety check: Notes | Apache-2.0 | today |
| 23 | A skill your agent uses when working with Slang shaders, shader modules, HLSL-compatible GPU code, graphics pipelines, compute shaders, tessellation, ray tracing, parameter blocks, generics… | github/ | 40k | 1 repo | ~1.8k | Automated safety check: Pass | MIT | today |
| 24 | How to build and run GPU targets under Buck in fbcode. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~998 | Automated safety check: Pass | MIT | today |
| 25 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 912 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 26 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.5k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | today |
| 27 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.5k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 28 | Continue KERMT pretraining on a custom SMILES corpus with a groverbase, cmim, or hybrid checkpoint. | NVIDIA/ | 3.5k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | today |
| 29 | Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint. | NVIDIA/ | 3.5k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 30 | Finetune a pretrained KERMT encoder on a labeled CSV. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | today |
| 31 | A skill your agent uses when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches. | NVIDIA/ | 3.5k | 1 repo | ~4.8k | Automated safety check: Pass | Apache-2.0 | today |
| 32 | Evaluate whether an existing hot path is a credible NVIDIA Warp candidate. | NVIDIA/ | 3.5k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 33 | A skill your agent uses when training, evaluating, or exporting Workflow policies with online RSL-RL or RLinf, including RL checkpoint and Workflow handoff. | NVIDIA/ | 3.5k | 1 repo | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 34 | Run predictions with a finetuned KERMT checkpoint on a SMILES-only CSV. | NVIDIA/ | 3.5k | 1 repo | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | Pretrain a fresh KERMT model from scratch on a user-provided corpus. | NVIDIA/ | 3.5k | 1 repo | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 36 | Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test… | NVIDIA/ | 3.5k | 1 repo | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 37 | Host setup for TAO GPU backends. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~3.4k | Automated safety check: Notes | Apache-2.0 | today |
| 38 | A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the… | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | today |
| 39 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.5k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | today |
| 40 | DALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks. | NVIDIA/ | 3.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | today |
| 41 | A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code. | NVIDIA/ | 3.5k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | today |
| 42 | A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance. | NVIDIA/ | 3.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 43 | Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI). | NVIDIA/ | 3.5k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | today |
| 44 | NVIDIA DeepStream SDK development with Python pyservicemaker API. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 45 | Build deterministic forecast scripts with Earth2Studio (model, data source, IO, inference). | NVIDIA/ | 3.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 46 | Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. | NVIDIA/ | 3.5k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | Used for generating synthetic body MRI volumes with NV-Generate-CTMR rflow-mr. | NVIDIA/ | 3.5k | — | ~2k | Automated safety check: Notes | Apache-2.0 | today |
| 48 | A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. | NVIDIA/ | 3.5k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
Questions, answered from the data.
What is the best CUDA skill?
Create Cuda Python Pull Request (official) from NVIDIA/cuda-python ranks first of the 280 CUDA skills listed here, with the highest score: its repository has 3.4k GitHub stars, its SKILL.md loads about 1.1k tokens and it passes the automated safety check with no findings. Next come CUTLASS FMHA Incremental Rebuild and ONNX Runtime Source Build.
Is there an official CUDA skill?
90 of the 280 CUDA skills are official, published by the vendor's own GitHub organization: Create Cuda Python Pull Request, CUTLASS FMHA Incremental Rebuild, ONNX Runtime Source Build, ONNX Runtime Release Notes, ONNX Runtime GPU Transformers Tests and 85 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.