Search
NVIDIA AI Platform · For developers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 242 | Deploy and operate the RTVI-CV-3D microservice as MV3DT (MODE=mv3dt): per-camera DeepStream perception plus BEV Fusion over calibrated cameras. | NVIDIA/ | 3.6k | — | ~4.8k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 243 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 244 | 244.Ollama Setup Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Notes | MIT | yesterday |
| 245 | Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment. | borghei/ | 891 | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 246 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 247 | Build Holoscan SDK from source via the in-tree ./run script. | NVIDIA/ | 3.6k | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 248 | Extract Jetson Linux + sample-rootfs tarballs and run applybinaries.sh for the active target, then record bspimage in the profile. | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 249 | Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 250 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 251 | Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 252 | Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~973 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 253 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 254 | MoE expert-parallel communication overlap in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 255 | Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 256 | Evidence-gated workflow for MoE performance optimization in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 257 | Practical guidance for training MoE VLMs in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 258 | Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~924 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 259 | A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS… | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 260 | Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Warn | Apache-2.0 | yesterday |
| 261 | Expert cuTile programming assistant. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 262 | Test system for Megatron-LM. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 263 | Deploy inference services on CoreWeave with Helm charts and Kustomize. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 264 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 265 | 265.Serving Systems LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 105 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 266 | Run GPU jobs on NVIDIA NIM microservices via host.compute.create('byoc:nvidia', ...). | PKU-YuanGroup/ | 622 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 267 | Verl 分布式训练服务一键拉起与配置。触发场景:(1) 用户要启动 Verl 训练任务或部署 RLHF/DAPO 训练环境 (2) 在 NPU 集群上拉起 Verl 训练容器 (3) 配置 Ray 集群和 SwanLab 监控 (4) 根据 7 位二进制掩码灵活配置加速特性。支持 Qwen3-8B 等 Megatron 模型的 DAPO 训练全流程。 | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | yesterday |
| 268 | Mod or remaster a game with RTX Remix - open and edit projects, swap textures and models. | NVIDIA/ | 3.6k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 269 | 269.GPU Optimizer GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 329 | — | ~3.5k | Automated safety check: Notes | MIT | 5 days ago |
| 270 | Use this sub-skill for Torch-TensorRT model compilation, dynamic input planning, torch.export workflows, save/load formats, raw TensorRT engines, and compile-time troubleshooting. | VectorSpaceLab/ | 331 | — | ~1.3k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 271 | Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark… | VectorSpaceLab/ | 331 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 272 | 272.Torch Tensorrt A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging… | VectorSpaceLab/ | 331 | — | ~1.5k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 273 | 273.Setup Prover Set up and deploy a Boundless prover to a GPU server using Ansible. | boundless-xyz/ | 193 | — | ~4.1k | Automated safety check: Warn | Apache-2.0 | 1 mo ago |
| 274 | MindSpeed-MM multimodal model suite environment setup guide for Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~3.1k | Automated safety check: Pass | No licence | yesterday |
| 275 | 275.Mindspeed Mm Vlm Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM. | ascend-ai-coding/ | 174 | — | ~5.2k | Automated safety check: Pass | No licence | yesterday |
| 276 | 276.Verl Quickstart Generates an executable, end-to-end VERL reinforcement learning quickstart runbook for Ascend/NPU (docker image, dataset preprocessing, model setup, mainppo training, and examples/run.sh flow). | ascend-ai-coding/ | 174 | — | ~592 | Automated safety check: Pass | No licence | yesterday |
| 277 | One-time session setup and orchestration map for the TAO skill bank. | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | yesterday |
| 278 | Linting and formatting for Megatron-LM. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~406 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 279 | 279.Parakeet Stt Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU). | sundial-org/ | 663 | — | ~771 | Automated safety check: Pass | No licence | 7 mo ago |
| 280 | 280.Cuda CUDA C/C++ skill for NVIDIA GPU kernel programming. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 252 | — | ~1.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 281 | 281.Cuda Debugging CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 252 | — | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 282 | 282.Cuda Profiling CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 252 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |
| 283 | 283.GPU Memory Model GPU memory model skill for SIMT execution and memory hierarchy. | mohitmishra786/ | 252 | — | ~1.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 284 | Bind pre-downloaded Jetson reference docs (developer guide, design guide, pinmux, schematics) into the active profile documents block. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 285 | Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar. | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | yesterday |
| 286 | The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective +… | NVIDIA/ | 3.6k | — | ~1.5k | Automated safety check: Warn | Apache-2.0 | yesterday |
| 287 | The Docker execution platform for TAO jobs — a local daemon or a remote GPU box via DOCKERHOST=ssh://user@host. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Warn | Apache-2.0 | yesterday |