Search
NVIDIA AI Platform · Container orchestration
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 2 | Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional… | NVIDIA/ | 159 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 3 | Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and… | NVIDIA/ | 159 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 4 | A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment. | NVIDIA/ | 1.9k | — | ~116 | Automated safety check: Pass | Apache-2.0 | 22 days ago |
| 5 | Multi-agent PR review using Claude Code, Codex, and CodeRabbit. | NVIDIA/ | 440 | — | ~15k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 7 | Start up, tear down, and configure the local Kubernetes development environment for OpenShell. | NVIDIA/ | 16k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster… | NVIDIA/ | 440 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | today |
| 9 | Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster | ai-runway/ | 102 | — | ~927 | Automated safety check: Pass | Apache-2.0 | 14 days ago |
| 10 | 10.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | 2 days ago |
| 11 | Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target. | microsoft/ | 126 | — | ~5.7k | Automated safety check: Notes | MIT | yesterday |
| 12 | A skill your agent uses when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which… | NVIDIA/ | 440 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 13 | A skill your agent uses when the user runs /aicr-triage or asks to triage, review, or clean up a GitHub org-level Projects v2 board (default NVIDIA AICR project 248). | NVIDIA/ | 440 | — | ~6.9k | Automated safety check: Pass | Apache-2.0 | today |
| 14 | Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. | sickn33/ | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | 2 days ago |
| 15 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 16 | Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes. | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 17 | Run container-backed AutoML / hyperparameter optimization (HPO) for NVIDIA TAO networks using AutoMLRunner. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 18 | Host setup for TAO GPU backends. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~3.4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 19 | Set up or troubleshoot the Auto Ontology runtime. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~2.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 20 | A skill your agent uses when the user is hands-on deploying an in-bundle DOCA service container (Argus, DMS, Firefly, or UROM service) on a BlueField — kubelet standalone watching a static-pod… | NVIDIA/ | 3.6k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 21 | The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action. | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 22 | A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS… | NVIDIA/ | 3.6k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 23 | Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Warn | Apache-2.0 | 2 days ago |
| 24 | Deploy inference services on CoreWeave with Helm charts and Kustomize. | jeremylongshore/ | 2.8k | — | ~1.5k | Automated safety check: Pass | MIT | yesterday |
| 25 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 26 | Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar. | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | 2 days ago |
| 27 | The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective +… | NVIDIA/ | 3.6k | — | ~1.5k | Automated safety check: Warn | Apache-2.0 | 2 days ago |