Search

NVIDIA AI Platform · Container orchestration

27 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

dstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.02 days ago
2

Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional…

NVIDIA/k8s-nim-operator159—~4.7kAutomated safety check: PassApache-2.05 days ago
3

Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and…

NVIDIA/k8s-nim-operator159—~3.6kAutomated safety check: PassApache-2.05 days ago
4

A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment.

NVIDIA/dcgm-exporter1.9k—~116Automated safety check: PassApache-2.022 days ago
5

Multi-agent PR review using Claude Code, Codex, and CodeRabbit.

NVIDIA/aicr440—~15kAutomated safety check: PassApache-2.0today
6

Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format.

dstackai/dstack2.3k—~403Automated safety check: PassMPL-2.02 days ago
7

Start up, tear down, and configure the local Kubernetes development environment for OpenShell.

NVIDIA/OpenShell16k—~4.9kAutomated safety check: PassApache-2.0yesterday
8

A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…

NVIDIA/aicr440—~3.5kAutomated safety check: PassApache-2.0today
9

Interactively build, push or load, and deploy an airunway component (controller or any provider) to the cluster

ai-runway/airunway102—~927Automated safety check: PassApache-2.014 days ago
10

dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

dstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.02 days ago
11

Set up a K3s cluster on an NVIDIA GPU host, connect it to Azure Arc, and configure Azure ML to use it as a Kubernetes compute target.

microsoft/physical-ai-toolchain126—~5.7kAutomated safety check: NotesMITyesterday
12

A skill your agent uses when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which…

NVIDIA/aicr440—~2.8kAutomated safety check: PassApache-2.0today
13
13.Aicr TriageOfficial

A skill your agent uses when the user runs /aicr-triage or asks to triage, review, or clean up a GitHub org-level Projects v2 board (default NVIDIA AICR project 248).

NVIDIA/aicr440—~6.9kAutomated safety check: PassApache-2.0today
14

Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.

sickn33/agentic-awesome-skills47k2 repos~3.2kAutomated safety check: PassMIT2 days ago
15

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

sickn33/agentic-awesome-skills47k1 repo~2.3kAutomated safety check: PassMIT2 days ago
16

Select, validate, patch, and deploy existing NVIDIA Dynamo Kubernetes recipes.

NVIDIA/skills3.6k—~1.8kAutomated safety check: PassApache-2.02 days ago
17
17.Tao Run AutomlOfficial

Run container-backed AutoML / hyperparameter optimization (HPO) for NVIDIA TAO networks using AutoMLRunner.

NVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.02 days ago
18

Host setup for TAO GPU backends. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~3.4kAutomated safety check: NotesApache-2.02 days ago
19

Set up or troubleshoot the Auto Ontology runtime. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k—~2.7kAutomated safety check: NotesApache-2.02 days ago
20

A skill your agent uses when the user is hands-on deploying an in-bundle DOCA service container (Argus, DMS, Firefly, or UROM service) on a BlueField — kubelet standalone watching a static-pod…

NVIDIA/skills3.6k—~2.5kAutomated safety check: PassApache-2.02 days ago
21

The mandatory pre-launch gate and four-verb execution contract for every TAO workflow or action.

NVIDIA/skills3.6k—~4.5kAutomated safety check: NotesApache-2.02 days ago
22

A skill your agent uses when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS…

NVIDIA/skills3.6k—~2.8kAutomated safety check: NotesApache-2.02 days ago
23

Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training.

NVIDIA/skills3.6k—~4.9kAutomated safety check: WarnApache-2.02 days ago
24

Deploy inference services on CoreWeave with Helm charts and Kustomize.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.5kAutomated safety check: PassMITyesterday
25

Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server.

BagelHole/DevOps-Security-Agent-Skills1.2k—~2.1kAutomated safety check: PassMIT4 mo ago
26

Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar.

NVIDIA/skills3.6k—~3.8kAutomated safety check: WarnApache-2.02 days ago
27
27.Tao Data IoOfficial

The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective +…

NVIDIA/skills3.6k—~1.5kAutomated safety check: WarnApache-2.02 days ago