Search

NVIDIA AI Platform · Deep learning

25 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document.

NVIDIA/Megatron-LM18k—~1.6kAutomated safety check: PassApache-2.0today
2

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
3

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
4

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k—~2kAutomated safety check: PassMIT5 days ago
5

Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
6

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark182—~1.4kAutomated safety check: PassMIT12 days ago
7

Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
8

Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
9

Guide for using SLIME (LLM post-training framework for RL Scaling).

yzlnew/infra-skills149—~3.2kAutomated safety check: PassNo licence3 mo ago
10

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0today
11

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT5 days ago
12

Run PyTorch training across GPUs with minimal changes. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1711 repo~2.3kAutomated safety check: PassMIT2 days ago
13

Official NVIDIA-authored guidance for navigating PhysicsNeMo — pick the model, datapipe, or example for a SciML/AI4Science task (surrogates, forecasting, downscaling, physics-informed, inverse…

NVIDIA/skills3.6k—~1.8kAutomated safety check: PassApache-2.0yesterday
14

Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and…

NVIDIA/skills3.6k—~3.5kAutomated safety check: PassApache-2.0yesterday
15

PyTorch-based TAO image classification. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~3.6kAutomated safety check: NotesApache-2.0yesterday
16

Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.0yesterday
17

Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA.

intel/torch-xpu-ops115—~681Automated safety check: PassApache-2.0today
18

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

majiayu000/spellbook287—~1.1kAutomated safety check: PassMIT2 days ago
19

Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.

NVIDIA/skills3.6k—~4.8kAutomated safety check: PassApache-2.0yesterday
20

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

NVIDIA/skills3.6k—~3.5kAutomated safety check: PassApache-2.0yesterday
21

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIA/skills3.6k—~973Automated safety check: PassApache-2.0yesterday
22

Practical guidance for training MoE VLMs in Megatron Bridge.

NVIDIA/skills3.6k—~1.3kAutomated safety check: PassApache-2.0yesterday
23

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

Mathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT4 days ago
24

A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging…

VectorSpaceLab/AREX-Skill331—~1.5kAutomated safety check: PassBSD-3-Clause1 mo ago
25

MindSpeed-MM multimodal model suite environment setup guide for Huawei Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~3.1kAutomated safety check: PassNo licencetoday