Language
CUDA agent skills, page 2
CUDA skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
Official
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | NV-Tesseract AD Diffusion — diffusion-based anomaly detection and fine-tuning for multivariate time series. | NVIDIA/ | 3.5k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | today |
| 50 | NV-Tesseract Forecasting — transformer-based multivariate time series forecasting with DARR (context-enhanced kNN retrieval), interpretability, and fine-tuning. | NVIDIA/ | 3.5k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | today |
| 51 | cuOpt REST server — start server, endpoints, Python/curl client examples. | NVIDIA/ | 3.5k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | today |
| 52 | Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store… | facebookexperimental/ | 201 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 53 | A skill your agent uses when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that… | NVIDIA/ | 3.5k | — | ~4k | Automated safety check: Pass | Apache-2.0 | today |
| 54 | A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation. | NVIDIA/ | 3.5k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | today |
| 55 | A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the… | NVIDIA/ | 3.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | today |
| 56 | A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests… | NVIDIA/ | 3.5k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 57 | A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under… | NVIDIA/ | 3.5k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 58 | Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment. | NVIDIA/ | 3.5k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | today |
| 59 | Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment. | NVIDIA/ | 3.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 60 | Install Holoscan SDK via the NGC Docker container. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 61 | Install Holoscan SDK natively on Ubuntu via apt. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | today |
| 62 | Install Holoscan SDK Python wheel via pip into a venv. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | today |
| 63 | Guides Holoscan SDK installation: inspects the host, assesses platform compatibility, recommends an install method, and delegates to the matching install skill. | NVIDIA/ | 3.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | today |
| 64 | Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | today |
| 65 | Used for command-shape or live NV-Reason-CXR chest X-ray reasoning smoke tests. | NVIDIA/ | 3.5k | — | ~3.9k | Automated safety check: Notes | Apache-2.0 | today |
| 66 | Used for running NV-Segment-CT VISTA3D on CT NIfTI volumes and recording label-map evidence. | NVIDIA/ | 3.5k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | today |
| 67 | Used for running NV-Segment-CTMR on CT or MRI NIfTI volumes and recording label-map evidence. | NVIDIA/ | 3.5k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | today |
| 68 | A skill your agent uses when authoring or validating PAIDF augmentation YAML configs, or running remote Cosmos Transfer (including Cosmos3 WSM controls), Cosmos Predict, image-edit, or… | NVIDIA/ | 3.5k | — | ~5.3k | Automated safety check: Pass | Apache-2.0 | today |
| 69 | How to create a pull request for the intel/torch-xpu-ops repository. | intel/ | 115 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 70 | Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads. | NVIDIA/ | 3.5k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 71 | A skill your agent uses when porting circuits from another framework (e.g. | NVIDIA/ | 3.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 72 | Install cuOpt for Python, C, or server via pip, conda, or Docker; verify the install. | NVIDIA/ | 3.5k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | today |
| 73 | Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. | NVIDIA/ | 3.5k | — | ~844 | Automated safety check: Pass | Apache-2.0 | today |
| 74 | NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. | NVIDIA/ | 3.5k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | today |
| 75 | How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT… | NVIDIA/ | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | today |
| 76 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 77 | Build Holoscan SDK from source via the in-tree ./run script. | NVIDIA/ | 3.5k | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | today |
| 78 | Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. | NVIDIA/ | 3.5k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 79 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 80 | Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 81 | Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.5k | — | ~973 | Automated safety check: Pass | Apache-2.0 | today |
| 82 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.5k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 83 | MoE expert-parallel communication overlap in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 84 | Representative, point-in-time MoE training playbooks by hardware and model family. | NVIDIA/ | 3.5k | — | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 85 | Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 86 | Evidence-gated workflow for MoE performance optimization in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 87 | Practical guidance for training MoE VLMs in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 88 | Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.5k | — | ~924 | Automated safety check: Pass | Apache-2.0 | today |
| 89 | One-time session setup and orchestration map for the TAO skill bank. | NVIDIA/ | 3.5k | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | today |
| 90 | The Docker execution platform for TAO jobs — a local daemon or a remote GPU box via DOCKERHOST=ssh://user@host. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Warn | Apache-2.0 | today |
Community
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 91 | Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang. | sgl-project/ | 37k | 3 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 92 | 92.Aoti Debug Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch. | pytorch/ | 104k | 1 repo | ~1.7k | Automated safety check: Pass | Unknown | today |
| 93 | 93.Paddle Build A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes. | PaddlePaddle/ | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 94 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 95 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 96 | Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm. | ztxz16/ | 5.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |