Search
NVIDIA AI Platform
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 385 | Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. | NVIDIA/ | 3.6k | — | ~844 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 386 | NVIDIA App MCP: drivers, games, laptops, overlay. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 387 | Coordinate the end-to-end CAD/source-asset to SimReady workflow. | NVIDIA/ | 3.6k | — | ~2.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 388 | Top-level workflow skill for USD performance diagnosis and optimization. | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 389 | NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 390 | Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality). | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 391 | Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. | NVIDIA/ | 3.6k | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 392 | How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT… | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 393 | How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 394 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 395 | Deploy and operate the RTVI-CV-3D microservice as MV3DT (MODE=mv3dt): per-camera DeepStream perception plus BEV Fusion over calibrated cameras. | NVIDIA/ | 3.6k | — | ~4.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 396 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 397 | 397.Ollama Setup Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Notes | MIT | yesterday |
| 398 | Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment. | borghei/ | 891 | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 399 | Optimize Rotary Position Embedding (RoPE) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~421 | Automated safety check: Pass | No licence | 6 mo ago |
| 400 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | 3 days ago |
| 401 | Build Holoscan SDK from source via the in-tree ./run script. | NVIDIA/ | 3.6k | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 402 | Bootstrap a custom carrier board by forking carrier files and scaffolding a DT overlay from the reference devkit. | NVIDIA/ | 3.6k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 403 | Build a per-target knowledge-base markdown next to the active profile by walking the BSP root and source tree. | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 404 | Extract Jetson Linux + sample-rootfs tarballs and run applybinaries.sh for the active target, then record bspimage in the profile. | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 405 | Entry skill for Jetson / IGX BSP customization. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 406 | Switch the active Jetson target-platform pointer to an existing profile YAML. | NVIDIA/ | 3.6k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 407 | Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings. | NVIDIA/ | 3.6k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 408 | Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. | NVIDIA/ | 3.6k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 409 | Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 410 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 411 | Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 412 | Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 413 | Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~973 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 414 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 415 | MoE expert-parallel communication overlap in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 416 | Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 417 | Evidence-gated workflow for MoE performance optimization in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 418 | Practical guidance for training MoE VLMs in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 419 | Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration. | NVIDIA/ | 3.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 420 | Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 421 | Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~924 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 422 | Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine. | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 423 | A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS… | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 424 | Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Warn | Apache-2.0 | 2 days ago |
| 425 | Expert cuTile programming assistant. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 426 | Test system for Megatron-LM. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 427 | Deploy inference services on CoreWeave with Helm charts and Kustomize. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 428 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 429 | 429.Serving Systems LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 105 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 430 | Run GPU jobs on NVIDIA NIM microservices via host.compute.create('byoc:nvidia', ...). | PKU-YuanGroup/ | 622 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 431 | Verl 分布式训练服务一键拉起与配置。触发场景:(1) 用户要启动 Verl 训练任务或部署 RLHF/DAPO 训练环境 (2) 在 NPU 集群上拉起 Verl 训练容器 (3) 配置 Ray 集群和 SwanLab 监控 (4) 根据 7 位二进制掩码灵活配置加速特性。支持 Qwen3-8B 等 Megatron 模型的 DAPO 训练全流程。 | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | yesterday |
| 432 | Mod or remaster a game with RTX Remix - open and edit projects, swap textures and models. | NVIDIA/ | 3.6k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |