Search
NVIDIA AI Platform · Deep learning
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 5 | Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 12 days ago |
| 7 | Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 8 | Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 9 | Guide for using SLIME (LLM post-training framework for RL Scaling). | yzlnew/ | 149 | — | ~3.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 10 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | today |
| 11 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 12 | 12.Accelerate Run PyTorch training across GPUs with minimal changes. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 13 | Official NVIDIA-authored guidance for navigating PhysicsNeMo — pick the model, datapipe, or example for a SciML/AI4Science task (surrogates, forecasting, downscaling, physics-informed, inverse… | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 14 | Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and… | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | PyTorch-based TAO image classification. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 16 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 17 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | today |
| 18 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | 2 days ago |
| 19 | Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings. | NVIDIA/ | 3.6k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 21 | Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~973 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 22 | Practical guidance for training MoE VLMs in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 23 | GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 329 | — | ~3.5k | Automated safety check: Notes | MIT | 4 days ago |
| 24 | A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging… | VectorSpaceLab/ | 331 | — | ~1.5k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 25 | MindSpeed-MM multimodal model suite environment setup guide for Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~3.1k | Automated safety check: Pass | No licence | today |