Language

CUDA agent skills, page 2

Skills #49–96 of 280, ranked by score.

CUDA skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Official

Official CUDA skills
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

NV-Tesseract AD Diffusion — diffusion-based anomaly detection and fine-tuning for multivariate time series.

NVIDIA/skills3.5k—~2.9kAutomated safety check: NotesApache-2.0today
50

NV-Tesseract Forecasting — transformer-based multivariate time series forecasting with DARR (context-enhanced kNN retrieval), interpretability, and fine-tuning.

NVIDIA/skills3.5k—~3.2kAutomated safety check: NotesApache-2.0today
51

cuOpt REST server — start server, endpoints, Python/curl client examples.

NVIDIA/skills3.5k—~1.5kAutomated safety check: PassApache-2.0today
52

Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…

facebookexperimental/triton201—~1.1kAutomated safety check: PassMITtoday
53

A skill your agent uses when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that…

NVIDIA/skills3.5k—~4kAutomated safety check: PassApache-2.0today
54
54.Doca GpiOfficial

A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation.

NVIDIA/skills3.5k—~3.9kAutomated safety check: PassApache-2.0today
55
55.Doca GpunetioOfficial

A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the…

NVIDIA/skills3.5k—~3.7kAutomated safety check: PassApache-2.0today
56

A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests…

NVIDIA/skills3.5k—~4.2kAutomated safety check: PassApache-2.0today
57

A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under…

NVIDIA/skills3.5k—~3.8kAutomated safety check: PassApache-2.0today
58

Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment.

NVIDIA/skills3.5k—~1.6kAutomated safety check: NotesApache-2.0today
59

Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment.

NVIDIA/skills3.5k—~2kAutomated safety check: PassApache-2.0today
60

Install Holoscan SDK via the NGC Docker container. An agent skill from NVIDIA/skills.

NVIDIA/skills3.5k—~1.9kAutomated safety check: PassApache-2.0today
61

Install Holoscan SDK natively on Ubuntu via apt. An agent skill from NVIDIA/skills.

NVIDIA/skills3.5k—~1.6kAutomated safety check: NotesApache-2.0today
62

Install Holoscan SDK Python wheel via pip into a venv. An agent skill from NVIDIA/skills.

NVIDIA/skills3.5k—~1.6kAutomated safety check: NotesApache-2.0today
63
63.Holoscan SetupOfficial

Guides Holoscan SDK installation: inspects the host, assesses platform compatibility, recommends an install method, and delegates to the matching install skill.

NVIDIA/skills3.5k—~2.5kAutomated safety check: PassApache-2.0today
64

Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.

NVIDIA/skills3.5k—~4.7kAutomated safety check: PassApache-2.0today
65
65.Nv Reason CxrOfficial

Used for command-shape or live NV-Reason-CXR chest X-ray reasoning smoke tests.

NVIDIA/skills3.5k—~3.9kAutomated safety check: NotesApache-2.0today
66
66.Nv Segment CtOfficial

Used for running NV-Segment-CT VISTA3D on CT NIfTI volumes and recording label-map evidence.

NVIDIA/skills3.5k—~2.1kAutomated safety check: NotesApache-2.0today
67
67.Nv Segment CtmrOfficial

Used for running NV-Segment-CTMR on CT or MRI NIfTI volumes and recording label-map evidence.

NVIDIA/skills3.5k—~2.3kAutomated safety check: NotesApache-2.0today
68

A skill your agent uses when authoring or validating PAIDF augmentation YAML configs, or running remote Cosmos Transfer (including Cosmos3 WSM controls), Cosmos Predict, image-edit, or…

NVIDIA/skills3.5k—~5.3kAutomated safety check: PassApache-2.0today
69

How to create a pull request for the intel/torch-xpu-ops repository.

intel/torch-xpu-ops115—~1.2kAutomated safety check: PassApache-2.0yesterday
70

Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.

NVIDIA/skills3.5k—~2.3kAutomated safety check: PassApache-2.0today
71
71.Cudaq ImportingOfficial

A skill your agent uses when porting circuits from another framework (e.g.

NVIDIA/skills3.5k—~1.9kAutomated safety check: PassApache-2.0today
72
72.Cuopt InstallOfficial

Install cuOpt for Python, C, or server via pip, conda, or Docker; verify the install.

NVIDIA/skills3.5k—~1.1kAutomated safety check: PassApache-2.0today
73

Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit.

NVIDIA/skills3.5k—~844Automated safety check: PassApache-2.0today
74
74.RAG BlueprintOfficial

NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage.

NVIDIA/skills3.5k—~2.8kAutomated safety check: NotesApache-2.0today
75

How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT…

NVIDIA/skills3.5k—~4.5kAutomated safety check: NotesApache-2.0today
76

Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA.

intel/torch-xpu-ops115—~681Automated safety check: PassApache-2.0yesterday
77

Build Holoscan SDK from source via the in-tree ./run script.

NVIDIA/skills3.5k—~1.5kAutomated safety check: NotesApache-2.0today
78

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

NVIDIA/skills3.5k—~2.3kAutomated safety check: PassApache-2.0today
79

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

NVIDIA/skills3.5k—~3.5kAutomated safety check: PassApache-2.0today
80

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.

NVIDIA/skills3.5k—~3.5kAutomated safety check: PassApache-2.0today
81

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIA/skills3.5k—~973Automated safety check: PassApache-2.0today
82

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM…

NVIDIA/skills3.5k—~3.6kAutomated safety check: PassApache-2.0today
83

MoE expert-parallel communication overlap in Megatron Bridge.

NVIDIA/skills3.5k—~1.9kAutomated safety check: PassApache-2.0today
84

Representative, point-in-time MoE training playbooks by hardware and model family.

NVIDIA/skills3.5k—~2kAutomated safety check: PassApache-2.0today
85

Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills.

NVIDIA/skills3.5k—~1.2kAutomated safety check: PassApache-2.0today
86

Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

NVIDIA/skills3.5k—~3.3kAutomated safety check: PassApache-2.0today
87

Practical guidance for training MoE VLMs in Megatron Bridge.

NVIDIA/skills3.5k—~1.3kAutomated safety check: PassApache-2.0today
88

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIA/skills3.5k—~924Automated safety check: PassApache-2.0today
89
89.Tao SetupOfficial

One-time session setup and orchestration map for the TAO skill bank.

NVIDIA/skills3.5k—~1.8kAutomated safety check: WarnApache-2.0today
90

The Docker execution platform for TAO jobs — a local daemon or a remote GPU box via DOCKERHOST=ssh://user@host.

NVIDIA/skills3.5k—~5kAutomated safety check: WarnApache-2.0today

Community

Community CUDA skills
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
91

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

sgl-project/sglang37k3 repos~2.1kAutomated safety check: PassApache-2.0today
92

Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

pytorch/pytorch104k1 repo~1.7kAutomated safety check: PassUnknowntoday
93

A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

PaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.07 days ago
94

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.07 days ago
95

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
96

Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

ztxz16/fastllm5.1k—~1.8kAutomated safety check: PassApache-2.0today