Search
AI & LLM Engineering · CUDA
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | Build deterministic forecast scripts with Earth2Studio (model, data source, IO, inference). | NVIDIA/ | 3.6k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 146 | Install or repair the FoundationPose perception pipeline and build its FoundationStereo TensorRT engines. | NVIDIA/ | 3.6k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 147 | Used for generating synthetic body MRI volumes with NV-Generate-CTMR rflow-mr. | NVIDIA/ | 3.6k | — | ~2k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 148 | A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. | NVIDIA/ | 3.6k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 149 | NV-Tesseract AD Diffusion — diffusion-based anomaly detection and fine-tuning for multivariate time series. | NVIDIA/ | 3.6k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 150 | NV-Tesseract Forecasting — transformer-based multivariate time series forecasting with DARR (context-enhanced kNN retrieval), interpretability, and fine-tuning. | NVIDIA/ | 3.6k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 151 | 151.Project Map Maps every HOT-Step CPP feature to its route file, service, UI folder, and engine subsystem, including port topology and the browser-to-engine request path. | scragnog/ | 174 | — | ~5.4k | Automated safety check: Notes | MIT | 2 days ago |
| 152 | Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent. | amd/ | 182 | — | ~4.8k | Automated safety check: Pass | MIT | 13 days ago |
| 153 | Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. | amd/ | 182 | — | ~1.9k | Automated safety check: Notes | MIT | 13 days ago |
| 154 | Install or verify the correct PyTorch build for a user's accelerator backend before Quark installation. | amd/ | 182 | — | ~1.6k | Automated safety check: Pass | MIT | 13 days ago |
| 155 | Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store… | facebookexperimental/ | 201 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 156 | Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. | matlab/ | 1.1k | — | ~2.8k | Automated safety check: Pass | Unknown | 3 days ago |
| 157 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 1.1k | — | ~4.6k | Automated safety check: Pass | Unknown | 3 days ago |
| 158 | A skill your agent uses when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that… | NVIDIA/ | 3.6k | — | ~4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 159 | A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation. | NVIDIA/ | 3.6k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 160 | A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the… | NVIDIA/ | 3.6k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 161 | A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests… | NVIDIA/ | 3.6k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 162 | A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under… | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 163 | Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment. | NVIDIA/ | 3.6k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 164 | Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment. | NVIDIA/ | 3.6k | — | ~2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 165 | Install Holoscan SDK natively on Ubuntu via apt. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 166 | Install Holoscan SDK Python wheel via pip into a venv. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 167 | Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 168 | Used for command-shape or live NV-Reason-CXR chest X-ray reasoning smoke tests. | NVIDIA/ | 3.6k | — | ~3.9k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 169 | Used for running NV-Segment-CT VISTA3D on CT NIfTI volumes and recording label-map evidence. | NVIDIA/ | 3.6k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 170 | Used for running NV-Segment-CTMR on CT or MRI NIfTI volumes and recording label-map evidence. | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 171 | Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. | NVIDIA/ | 3.6k | — | ~844 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 172 | NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 173 | How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT… | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 174 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 175 | 175.Mat Lammps Md Build and run LAMMPS molecular dynamics with isolated MLIP-specific binaries (MACE, MatGL/CHGNet, FairChem) to avoid Python and Torch stack conflicts. | learningmatter-mit/ | 176 | — | ~1.2k | Automated safety check: Pass | MIT | 3 days ago |
| 176 | Generate inorganic material structures using MatterGen, a diffusion-based generative model. | learningmatter-mit/ | 176 | — | ~1.8k | Automated safety check: Pass | MIT | 3 days ago |
| 177 | Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 178 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 179 | Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 180 | Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~973 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 181 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 182 | MoE expert-parallel communication overlap in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 183 | Representative, point-in-time MoE training playbooks by hardware and model family. | NVIDIA/ | 3.6k | — | ~2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 184 | Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 185 | Practical guidance for training MoE VLMs in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 186 | Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~924 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 187 | 187.Flash Attention Optimize transformer attention with Flash Attention — 2-4x speedup, 10-20x memory reduction for long sequences on CUDA GPUs. | AlexAI-MCP/ | 135 | — | ~1.2k | Automated safety check: Pass | MIT | 6 mo ago |
| 188 | Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 189 | 189.GPU Optimizer GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 329 | — | ~3.5k | Automated safety check: Notes | MIT | 5 days ago |
| 190 | 190.Hunyuan Video Use this operating skill for Tencent-Hunyuan/HunyuanVideo text-to-video setup, checkpoint layout, inference commands, Gradio launch, and CUDA/FP8/xDiT troubleshooting. | VectorSpaceLab/ | 331 | — | ~1.1k | Automated safety check: Pass | Unknown | 1 mo ago |
| 191 | 191.Make It 3D Use this repo skill for Make-It-3D single-image 3D creation, including CUDA asset setup, alpha-image validation, coarse NeRF optimization, refinement, rendering, export, and troubleshooting. | VectorSpaceLab/ | 331 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 192 | Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark… | VectorSpaceLab/ | 331 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |