Search
DevOps & Cloud · NVIDIA AI Platform · For devops and sre engineers
- DevOps & Cloud (remove filter)
- NVIDIA AI Platform (remove filter)
- For devops and sre engineers (remove filter)
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up. | NVIDIA/ | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Optimize AlpaSim Slurm topology throughput using persistent local Prometheus/Grafana telemetry and run artifacts. | NVlabs/ | 1.3k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 21 days ago |
| 4 | Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional… | NVIDIA/ | 159 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 5 | Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and… | NVIDIA/ | 159 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 6 | A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment. | NVIDIA/ | 1.9k | — | ~116 | Automated safety check: Pass | Apache-2.0 | 21 days ago |
| 7 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 456 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a. | brevdev/ | 146 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 10 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | today |
| 11 | Prepare, install and commission NVIDIA GPU hosts for Vast.ai using host-specific safety gates, XFS storage checks, cgroup/runtime testing and optional VM qualification. | jjziets/ | 162 | — | ~2.1k | Automated safety check: Pass | No licence | 6 days ago |
| 12 | Create safe, privacy-preserving Windows 11 C-drive cleanup, storage migration, cache relocation, package-manager policy, and maintenance plans. | Nongfsq/ | 154 | — | ~845 | Automated safety check: Pass | MIT | 4 mo ago |
| 13 | Use BEFORE running a full CompileIQ search. An agent skill from NVIDIA/CompileIQ. | NVIDIA/ | 138 | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | 16 days ago |
| 14 | Start up, tear down, and configure the local Kubernetes development environment for OpenShell. | NVIDIA/ | 16k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 15 | A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster… | NVIDIA/ | 440 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 16 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster | ai-runway/ | 102 | — | ~927 | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 18 | A skill your agent uses when migrating applications, examples, integrations, documentation, manifests, or repository code from NeMo Flow to NeMo Relay across Python, Rust, Node.js, Go, C FFI, CLI… | NVIDIA/ | 192 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 19 | 19.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | today |
| 20 | Best practices for Docker-based ROS2 development including multi-stage Dockerfiles, docker-compose for multi-container robotic systems, DDS discovery across containers, GPU passthrough for… | arpitg1304/ | 369 | — | ~9.1k | Automated safety check: Notes | Apache-2.0 | 1 mo ago |
| 21 | A skill your agent uses when asked to install, deploy, run, validate, troubleshoot, or stop NVIDIA Deep Researcher Agent Blueprint infrastructure. | NVIDIA-AI-Blueprints/ | 885 | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 22 | Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification. | drawthingsai/ | 582 | — | ~2.2k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 23 | Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. | microsoft/ | 123 | — | ~5.7k | Automated safety check: Notes | MIT | yesterday |
| 24 | A skill your agent uses when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which… | NVIDIA/ | 440 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | 25.Upgrade Deps Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL. | areal-project/ | 5.8k | — | ~6k | Automated safety check: Pass | Apache-2.0 | today |
| 26 | Guidance for Azure Confidential Computing — protecting data in use through hardware-based Trusted Execution Environments (TEEs). | vinayaklatthe/ | 175 | — | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. | sickn33/ | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | today |
| 28 | Generate and analyze DNA sequences using NVIDIA's Evo 2 BioNeMo NIM microservice. | NVIDIA/ | 3.5k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 29 | Generate multiple sequence alignments (MSAs) for protein sequences using the ColabFold MSA-Search NIM. | NVIDIA/ | 3.5k | 1 repo | ~4.6k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 30 | Verify whether a tagged NeMo Relay release reached GitHub Actions, tagged Go module source, crates.io, PyPI, and npm. | NVIDIA/ | 192 | — | ~415 | Automated safety check: Pass | Apache-2.0 | today |
| 31 | Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. | NVIDIA/ | 3.5k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | Generate novel drug-like molecules using the GenMol NIM microservice. | NVIDIA/ | 3.5k | 1 repo | ~1.4k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 33 | A skill your agent uses for MolMIM, NVIDIA's BioNeMo NIM microservice for small-molecule latent-space generation and optimization. | NVIDIA/ | 3.5k | 1 repo | ~1.9k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 34 | A skill your agent uses for OpenFold2, NVIDIA's BioNeMo NIM microservice for monomer protein structure prediction. | NVIDIA/ | 3.5k | 1 repo | ~1.8k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 35 | A skill your agent uses for OpenFold3, NVIDIA's BioNeMo NIM microservice for biomolecular structure prediction. | NVIDIA/ | 3.5k | 1 repo | ~1.9k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 36 | Run RFDiffusion protein backbone design via NVIDIA NIM. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | 1 repo | ~1.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 37 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | today |
| 38 | Bootstrap the KERMT agent environment — verify host docker + nvidia-container-toolkit, build the kermt:latest image from the repo's Dockerfile if it doesn't yet exist, and run a GPU smoke test… | NVIDIA/ | 3.5k | 1 repo | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 39 | Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes. | NVIDIA/ | 3.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 40 | Run container-backed AutoML / hyperparameter optimization (HPO) for NVIDIA TAO networks using AutoMLRunner. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 41 | Host setup for TAO GPU backends. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~3.4k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 42 | A skill your agent uses for VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, Slack notifications, incident queries, camera onboarding. | NVIDIA/ | 3.5k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 43 | Use Boltz2 NIM for biomolecular structure prediction and binding affinity. | NVIDIA/ | 3.5k | 1 repo | ~1.4k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 44 | 44.Slime RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 158 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | yesterday |
| 45 | Model and publish semantic definitions in Auto Ontology. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 46 | Set up or troubleshoot the Auto Ontology runtime. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~2.7k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 47 | A skill your agent uses to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon… | NVIDIA/ | 3.5k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 48 | A skill your agent uses when the user is hands-on deploying an in-bundle DOCA service container (Argus, DMS, Firefly, or UROM service) on a BlueField — kubelet standalone watching a static-pod… | NVIDIA/ | 3.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | yesterday |