Search
AI & LLM Engineering · NVIDIA AI Platform
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 193 | Stage 3 of Clinical ASR Flywheel. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 194 | Enable MIPI/GMSL camera sensors on a Jetson Thor or Orin custom carrier by rendering a kernel-DT overlay from the in-tree sensor DTSI. | NVIDIA/ | 3.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 195 | Per-controller PCIe enable / disable / lanes / link-speed for a Jetson Thor or Orin custom carrier via ODMDATA + kernel-DT overlay. | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 196 | Configure Jetson UPHY lane allocation (uphy0/uphy1-config) on Orin/Thor custom carriers. | NVIDIA/ | 3.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 197 | Enable/disable Jetson USB2/USB3 SS ports via kernel-DT overlay. | NVIDIA/ | 3.6k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 198 | Set up the BSP source workspace: LinuxforTegra overlay tracker, bspsources, Crosstool-NG toolchain. | NVIDIA/ | 3.6k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 199 | Author a new Jetson target-platform profile (referencedevkit + optional customcarrier) and update the active pointer. | NVIDIA/ | 3.6k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 200 | Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. | NVIDIA/ | 3.6k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 201 | Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. | NVIDIA/ | 3.6k | — | ~844 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 202 | NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 203 | Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality). | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 204 | Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. | NVIDIA/ | 3.6k | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 205 | How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT… | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 206 | How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 207 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 208 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 209 | 209.Ollama Setup Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs. | jeremylongshore/ | 2.8k | — | ~1.4k | Automated safety check: Notes | MIT | yesterday |
| 210 | Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment. | borghei/ | 891 | — | ~1.8k | Automated safety check: Pass | MIT | 4 days ago |
| 211 | Optimize Rotary Position Embedding (RoPE) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~421 | Automated safety check: Pass | No licence | 6 mo ago |
| 212 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | 3 days ago |
| 213 | Bootstrap a custom carrier board by forking carrier files and scaffolding a DT overlay from the reference devkit. | NVIDIA/ | 3.6k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 214 | Build a per-target knowledge-base markdown next to the active profile by walking the BSP root and source tree. | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 215 | Extract Jetson Linux + sample-rootfs tarballs and run applybinaries.sh for the active target, then record bspimage in the profile. | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 216 | Entry skill for Jetson / IGX BSP customization. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 217 | Switch the active Jetson target-platform pointer to an existing profile YAML. | NVIDIA/ | 3.6k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 218 | Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings. | NVIDIA/ | 3.6k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 219 | Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. | NVIDIA/ | 3.6k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 220 | Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. | NVIDIA/ | 3.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 221 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 222 | Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 223 | Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~973 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 224 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 225 | MoE expert-parallel communication overlap in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 226 | Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 227 | Practical guidance for training MoE VLMs in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 228 | Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.6k | — | ~924 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 229 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 230 | 230.Serving Systems LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys. | uw-syfi/ | 105 | — | ~2.9k | Automated safety check: Pass | MIT | yesterday |
| 231 | Verl 分布式训练服务一键拉起与配置。触发场景:(1) 用户要启动 Verl 训练任务或部署 RLHF/DAPO 训练环境 (2) 在 NPU 集群上拉起 Verl 训练容器 (3) 配置 Ray 集群和 SwanLab 监控 (4) 根据 7 位二进制掩码灵活配置加速特性。支持 Qwen3-8B 等 Megatron 模型的 DAPO 训练全流程。 | ascend-ai-coding/ | 174 | — | ~2k | Automated safety check: Pass | No licence | yesterday |
| 232 | 232.GPU Optimizer GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 329 | — | ~3.5k | Automated safety check: Notes | MIT | 5 days ago |
| 233 | Use this sub-skill for Torch-TensorRT model compilation, dynamic input planning, torch.export workflows, save/load formats, raw TensorRT engines, and compile-time troubleshooting. | VectorSpaceLab/ | 331 | — | ~1.3k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 234 | Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark… | VectorSpaceLab/ | 331 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 235 | 235.Torch Tensorrt A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging… | VectorSpaceLab/ | 331 | — | ~1.5k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 236 | Generate migration deliverables for bringing relevant Megatron changes into MindSpeed after branch alignment and impact mapping are complete. | ascend-ai-coding/ | 174 | — | ~1.3k | Automated safety check: Pass | No licence | yesterday |
| 237 | Verl 单异步 DAPO 训练配置生成器。触发场景:(1) 启动单异步 DAPO 训练 (2) 生成训练脚本 (3) 配置特性参数 (4) 训练前检查。特性策略:用户未指定时默认开启性能特性(flashattn/dynamicbatch/removepadding/gradientcheckpointing),显存特性(offload/recompute)默认关闭。OOM… | ascend-ai-coding/ | 174 | — | ~1.2k | Automated safety check: Pass | No licence | yesterday |
| 238 | MindSpeed-MM multimodal model suite environment setup guide for Huawei Ascend NPU. | ascend-ai-coding/ | 174 | — | ~3.1k | Automated safety check: Pass | No licence | yesterday |
| 239 | 239.Mindspeed Mm Vlm Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM. | ascend-ai-coding/ | 174 | — | ~5.2k | Automated safety check: Pass | No licence | yesterday |
| 240 | 240.Verl Quickstart Generates an executable, end-to-end VERL reinforcement learning quickstart runbook for Ascend/NPU (docker image, dataset preprocessing, model setup, mainppo training, and examples/run.sh flow). | ascend-ai-coding/ | 174 | — | ~592 | Automated safety check: Pass | No licence | yesterday |