Search
DevOps & Cloud · NVIDIA AI Platform
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI. | NVIDIA/ | 18k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 2 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 4 days ago |
| 3 | Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up. | NVIDIA/ | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 4 | Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. | NVlabs/ | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 20 days ago |
| 5 | Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional… | NVIDIA/ | 159 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 6 | Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~15k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 7 | Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and… | NVIDIA/ | 159 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 8 | A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment. | NVIDIA/ | 1.9k | — | ~116 | Automated safety check: Pass | Apache-2.0 | 20 days ago |
| 9 | Multi-agent PR review using Claude Code, Codex, and CodeRabbit. | NVIDIA/ | 439 | — | ~15k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 443 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a. | brevdev/ | 144 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 12 | Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 13 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | yesterday |
| 14 | Prepare, install and commission NVIDIA GPU hosts for Vast.ai using host-specific safety gates, XFS storage checks, cgroup/runtime testing and optional VM qualification. | jjziets/ | 162 | — | ~2.1k | Automated safety check: Pass | No licence | 5 days ago |
| 15 | Prepare, validate, build, and use Nemotron Customizer airgap image bundles for offline clusters. | NVIDIA-NeMo/ | 2.1k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 16 | Create safe, privacy-preserving Windows 11 C-drive cleanup, storage migration, cache relocation, package-manager policy, and maintenance plans. | Nongfsq/ | 144 | — | ~845 | Automated safety check: Pass | MIT | 4 mo ago |
| 17 | Use BEFORE running a full CompileIQ search. An agent skill from NVIDIA/CompileIQ. | NVIDIA/ | 137 | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | 15 days ago |
| 18 | Start up, tear down, and configure the local Kubernetes development environment for OpenShell. | NVIDIA/ | 15k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 19 | A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster… | NVIDIA/ | 439 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 21 | 21.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 146 | — | ~547 | Automated safety check: Pass | No licence | 9 days ago |
| 22 | Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster | ai-runway/ | 101 | — | ~927 | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 23 | A skill your agent uses when migrating applications, examples, integrations, documentation, manifests, or repository code from NeMo Flow to NeMo Relay across Python, Rust, Node.js, Go, C FFI, CLI… | NVIDIA/ | 192 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 24 | 24.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | yesterday |
| 25 | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. | Orchestra-Research/ | 13k | 3 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 26 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 27 | Best practices for Docker-based ROS2 development including multi-stage Dockerfiles, docker-compose for multi-container robotic systems, DDS discovery across containers, GPU passthrough for… | arpitg1304/ | 368 | — | ~9.1k | Automated safety check: Notes | Apache-2.0 | 1 mo ago |
| 28 | A skill your agent uses when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA Deep Researcher Agent Blueprint infrastructure. | NVIDIA-AI-Blueprints/ | 883 | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 29 | Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification. | drawthingsai/ | 580 | — | ~2.2k | Automated safety check: Pass | GPL-3.0 | 2 days ago |
| 30 | Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. | microsoft/ | 123 | — | ~5.7k | Automated safety check: Notes | MIT | yesterday |
| 31 | A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~5.1k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 32 | Run, monitor, stop and report Primus convergence tests -- training a model on a real corpus and checking that the loss curve is healthy -- from a plain-language request such as "run convergence test… | AMD-AGI/ | 131 | — | ~2.1k | Automated safety check: Pass | Unknown | yesterday |
| 33 | A skill your agent uses when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which… | NVIDIA/ | 439 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 34 | 为 AutoResearch 的双 Agent 轨迹、付费 GPU 长跑、Docker 执行、可信评测与恢复建立共享协议、成本决策和隔离边界。用于小时/包日选择、启动或恢复 campaign、设计证据与防止题目或轨迹串用;不替代具体任务算法或最终平台 QA。 | bosprimigenious/ | 151 | — | ~553 | Automated safety check: Pass | MIT | 4 days ago |
| 35 | 35.Upgrade Deps Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL. | areal-project/ | 5.8k | — | ~6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 36 | Classify one failed NemoClaw GitHub Actions job using bounded, redacted logs and optional retained artifacts. | NVIDIA/ | 23k | — | ~806 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 37 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 103 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 38 | Guidance for Azure Confidential Computing — protecting data in use through hardware-based Trusted Execution Environments (TEEs). | vinayaklatthe/ | 175 | — | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 39 | Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. | sickn33/ | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 40 | Set up and manage NVIDIA GPU servers for AI workloads. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~2k | Automated safety check: Notes | MIT | yesterday |
| 41 | Generate and analyze DNA sequences using NVIDIA's Evo 2 BioNeMo NIM microservice. | NVIDIA/ | 3.5k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 42 | Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. | NVIDIA/ | 3.5k | 1 repo | ~4.6k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 43 | Run ProteinMPNN inverse folding via NVIDIA NIM to design protein sequences for a target backbone. | NVIDIA/ | 3.5k | 1 repo | ~2k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 44 | A skill your agent uses for VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, Slack notifications, incident queries, camera onboarding. | NVIDIA/ | 3.5k | 1 repo | ~4.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 45 | Verify whether a tagged NeMo Relay release reached GitHub Actions, tagged Go module source, crates.io, PyPI, and npm. | NVIDIA/ | 192 | — | ~415 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 46 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 47 | Run DiffDock molecular docking via NVIDIA NIM to predict small-molecule binding poses against protein targets. | NVIDIA/ | 3.5k | 1 repo | ~1.1k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 48 | Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. | NVIDIA/ | 3.5k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |