AI model or service
NVIDIA AI Platform agent skills, page 6
NVIDIA AI Platform skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
Official
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 241 | Top-level workflow skill for USD performance diagnosis and optimization. | NVIDIA/ | 3.5k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 242 | NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. | NVIDIA/ | 3.5k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | today |
| 243 | Filesystem RAG benchmarks: corpus/, train.json, evaluaterag.py (RAGAS quality). | NVIDIA/ | 3.5k | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | today |
| 244 | How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT… | NVIDIA/ | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | today |
| 245 | How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | today |
| 246 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | today |
| 247 | Deploy and operate the RTVI-CV-3D microservice as MV3DT (MODE=mv3dt): per-camera DeepStream perception plus BEV Fusion over calibrated cameras. | NVIDIA/ | 3.5k | — | ~4.8k | Automated safety check: Notes | Apache-2.0 | today |
| 248 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | today |
| 249 | Build Holoscan SDK from source via the in-tree ./run script. | NVIDIA/ | 3.5k | — | ~1.5k | Automated safety check: Notes | Apache-2.0 | today |
| 250 | Bootstrap a custom carrier board by forking carrier files and scaffolding a DT overlay from the reference devkit. | NVIDIA/ | 3.5k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 251 | Build a per-target knowledge-base markdown next to the active profile by walking the BSP root and source tree. | NVIDIA/ | 3.5k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 252 | Entry skill for Jetson / IGX BSP customization. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | today |
| 253 | Switch the active Jetson target-platform pointer to an existing profile YAML. | NVIDIA/ | 3.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 254 | Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings. | NVIDIA/ | 3.5k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | today |
| 255 | Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. | NVIDIA/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 256 | Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. | NVIDIA/ | 3.5k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 257 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 258 | Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 259 | Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.5k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 260 | Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.5k | — | ~973 | Automated safety check: Pass | Apache-2.0 | today |
| 261 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.5k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 262 | MoE expert-parallel communication overlap in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 263 | Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 264 | Evidence-gated workflow for MoE performance optimization in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 265 | Practical guidance for training MoE VLMs in Megatron Bridge. | NVIDIA/ | 3.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 266 | Operational guide for choosing and combining parallelism strategies in Megatron Bridge, including sizing rules, hardware topology mapping, and combined parallelism configuration. | NVIDIA/ | 3.5k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 267 | Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 268 | Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. | NVIDIA/ | 3.5k | — | ~924 | Automated safety check: Pass | Apache-2.0 | today |
| 269 | Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine. | NVIDIA/ | 3.5k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 270 | A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS… | NVIDIA/ | 3.5k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | today |
| 271 | Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training. | NVIDIA/ | 3.5k | — | ~4.9k | Automated safety check: Warn | Apache-2.0 | today |
| 272 | Expert cuTile programming assistant. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Pass | Apache-2.0 | today |
| 273 | Mod or remaster a game with RTX Remix - open and edit projects, swap textures and models. | NVIDIA/ | 3.5k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 274 | One-time session setup and orchestration map for the TAO skill bank. | NVIDIA/ | 3.5k | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | today |
| 275 | Bind pre-downloaded Jetson reference docs (developer guide, design guide, pinmux, schematics) into the active profile documents block. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 276 | Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar. | NVIDIA/ | 3.5k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | today |
Community
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 277 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 278 | Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search. | decolua/ | 30k | — | ~604 | Automated safety check: Pass | MIT | 7 days ago |
| 279 | Shows how to call NEAR AI Cloud through an OpenAI-compatible API and verify that inference ran in a TEE, using attestation checks and signed chat responses. | internet-court/ | 6.4k | 2 repos | ~1.3k | Automated safety check: Pass | Unknown | 1 mo ago |
| 280 | Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others. | decolua/ | 30k | — | ~745 | Automated safety check: Pass | MIT | 7 days ago |
| 281 | 281.Doc Reviewer Reviews recent code changes and checks if documentation needs updates. | NVlabs/ | 1.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 20 days ago |
| 282 | 282.Refactor Op Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code). | CVCUDA/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Unknown | 21 days ago |
| 283 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | today |
| 284 | Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA… | fla-org/ | 5.8k | — | ~4.2k | Automated safety check: Pass | MIT | today |
| 285 | 285.Local OCR Extract text from local images (PNG, JPEG) with precise coordinates, confidence scores, and stable error handling using the light-ocr CLI. | arcships/ | 610 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | today |
| 286 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 3 days ago |
| 287 | Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided… | NVIDIA-BioNeMo/ | 478 | — | ~3.1k | Automated safety check: Notes | Apache-2.0 | today |
| 288 | 288.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 10 days ago |