Search

NVIDIA AI Platform · For developers

287 skills found, page 6.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
241

Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline).

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.0yesterday
242

Deploy and operate the RTVI-CV-3D microservice as MV3DT (MODE=mv3dt): per-camera DeepStream perception plus BEV Fusion over calibrated cameras.

NVIDIA/skills3.6k—~4.8kAutomated safety check: NotesApache-2.0yesterday
243
243.Oob Perf AnalysisOfficial

Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA.

intel/torch-xpu-ops115—~681Automated safety check: PassApache-2.0yesterday
244

Configure auto-configure Ollama when user needs local LLM deployment, free AI alternatives, or wants to eliminate hosted API costs.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.4kAutomated safety check: NotesMITyesterday
245

Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment.

borghei/Claude-Skills891—~1.8kAutomated safety check: PassMIT4 days ago
246

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

majiayu000/spellbook287—~1.1kAutomated safety check: PassMIT2 days ago
247

Build Holoscan SDK from source via the in-tree ./run script.

NVIDIA/skills3.6k—~1.5kAutomated safety check: NotesApache-2.0yesterday
248
248.Jetson Init ImageOfficial

Extract Jetson Linux + sample-rootfs tarballs and run applybinaries.sh for the active target, then record bspimage in the profile.

NVIDIA/skills3.6k—~1.8kAutomated safety check: NotesApache-2.0yesterday
249

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

NVIDIA/skills3.6k—~2.3kAutomated safety check: PassApache-2.0yesterday
250

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

NVIDIA/skills3.6k—~3.5kAutomated safety check: PassApache-2.0yesterday
251

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.

NVIDIA/skills3.6k—~3.5kAutomated safety check: PassApache-2.0yesterday
252

Operational guide for enabling Megatron FSDP in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIA/skills3.6k—~973Automated safety check: PassApache-2.0yesterday
253

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM…

NVIDIA/skills3.6k—~3.6kAutomated safety check: PassApache-2.0yesterday
254

MoE expert-parallel communication overlap in Megatron Bridge.

NVIDIA/skills3.6k—~1.9kAutomated safety check: PassApache-2.0yesterday
255

Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~1.2kAutomated safety check: PassApache-2.0yesterday
256

Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.0yesterday
257

Practical guidance for training MoE VLMs in Megatron Bridge.

NVIDIA/skills3.6k—~1.3kAutomated safety check: PassApache-2.0yesterday
258

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

NVIDIA/skills3.6k—~924Automated safety check: PassApache-2.0yesterday
259

A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS…

NVIDIA/skills3.6k—~2.8kAutomated safety check: NotesApache-2.0yesterday
260

Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training.

NVIDIA/skills3.6k—~4.9kAutomated safety check: WarnApache-2.0yesterday
261

Expert cuTile programming assistant. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~5kAutomated safety check: PassApache-2.0yesterday
262
262.Mcore TestingOfficial

Test system for Megatron-LM. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~1.8kAutomated safety check: PassApache-2.0yesterday
263

Deploy inference services on CoreWeave with Helm charts and Kustomize.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.5kAutomated safety check: PassMITyesterday
264

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2.1kAutomated safety check: PassMIT4 mo ago
265

LLM and multimodal serving systems. An agent skill from uw-syfi/vibesys.

uw-syfi/vibesys105—~2.9kAutomated safety check: PassMITyesterday
266

Run GPU jobs on NVIDIA NIM microservices via host.compute.create('byoc:nvidia', ...).

PKU-YuanGroup/OpenAI4S622—~2.9kAutomated safety check: PassApache-2.02 days ago
267

Verl 分布式训练服务一键拉起与配置。触发场景:(1) 用户要启动 Verl 训练任务或部署 RLHF/DAPO 训练环境 (2) 在 NPU 集群上拉起 Verl 训练容器 (3) 配置 Ray 集群和 SwanLab 监控 (4) 根据 7 位二进制掩码灵活配置加速特性。支持 Qwen3-8B 等 Megatron 模型的 DAPO 训练全流程。

ascend-ai-coding/awesome-ascend-skills174—~2kAutomated safety check: PassNo licenceyesterday
268
268.Rtx Remix ModdingOfficial

Mod or remaster a game with RTX Remix - open and edit projects, swap textures and models.

NVIDIA/skills3.6k—~2.2kAutomated safety check: PassApache-2.0yesterday
269

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

Mathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT5 days ago
270

Use this sub-skill for Torch-TensorRT model compilation, dynamic input planning, torch.export workflows, save/load formats, raw TensorRT engines, and compile-time troubleshooting.

VectorSpaceLab/AREX-Skill331—~1.3kAutomated safety check: PassBSD-3-Clause1 mo ago
271

Use this sub-skill for Torch-TensorRT runtime performance controls, CUDA Graphs, output allocation, caches, TensorRT-RTX runtime settings, mutable modules, refit, weight streaming, and benchmark…

VectorSpaceLab/AREX-Skill331—~1kAutomated safety check: PassBSD-3-Clause1 mo ago
272

A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging…

VectorSpaceLab/AREX-Skill331—~1.5kAutomated safety check: PassBSD-3-Clause1 mo ago
273

Set up and deploy a Boundless prover to a GPU server using Ansible.

boundless-xyz/boundless193—~4.1kAutomated safety check: WarnApache-2.01 mo ago
274

MindSpeed-MM multimodal model suite environment setup guide for Huawei Ascend NPU.

ascend-ai-coding/awesome-ascend-skills174—~3.1kAutomated safety check: PassNo licenceyesterday
275

Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM.

ascend-ai-coding/awesome-ascend-skills174—~5.2kAutomated safety check: PassNo licenceyesterday
276

Generates an executable, end-to-end VERL reinforcement learning quickstart runbook for Ascend/NPU (docker image, dataset preprocessing, model setup, mainppo training, and examples/run.sh flow).

ascend-ai-coding/awesome-ascend-skills174—~592Automated safety check: PassNo licenceyesterday
277
277.Tao SetupOfficial

One-time session setup and orchestration map for the TAO skill bank.

NVIDIA/skills3.6k—~1.8kAutomated safety check: WarnApache-2.0yesterday
278

Linting and formatting for Megatron-LM. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~406Automated safety check: PassApache-2.0yesterday
279

Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU).

sundial-org/awesome-openclaw-skills663—~771Automated safety check: PassNo licence7 mo ago
280
280.Cuda

CUDA C/C++ skill for NVIDIA GPU kernel programming. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.9kAutomated safety check: PassMIT3 mo ago
281

CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.5kAutomated safety check: PassMIT3 mo ago
282

CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.6kAutomated safety check: NotesMIT3 mo ago
283

GPU memory model skill for SIMT execution and memory hierarchy.

mohitmishra786/low-level-dev-skills252—~1.9kAutomated safety check: PassMIT3 mo ago
284
284.Jetson Link DocsOfficial

Bind pre-downloaded Jetson reference docs (developer guide, design guide, pinmux, schematics) into the active profile documents block.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.0yesterday
285

Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar.

NVIDIA/skills3.6k—~3.8kAutomated safety check: WarnApache-2.0yesterday
286
286.Tao Data IoOfficial

The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective +…

NVIDIA/skills3.6k—~1.5kAutomated safety check: WarnApache-2.0yesterday
287
287.Tao Run On DockerOfficial

The Docker execution platform for TAO jobs — a local daemon or a remote GPU box via DOCKERHOST=ssh://user@host.

NVIDIA/skills3.6k—~5kAutomated safety check: WarnApache-2.0yesterday