Aicr Analyzing Snapshots
NVIDIA/aicr
A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…
Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-run-on-kubernetes --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-run-on-kubernetes .claude/skills/tao-run-on-kubernetes && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-run-on-kubernetes" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-kubernetes into .claude/skills/tao-run-on-kubernetes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-kubernetes", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-kubernetesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-run-on-kubernetes --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-run-on-kubernetes .agents/skills/tao-run-on-kubernetes && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-run-on-kubernetes" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-kubernetes into .agents/skills/tao-run-on-kubernetes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-kubernetes", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-run-on-kubernetes --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-run-on-kubernetes .cursor/skills/tao-run-on-kubernetes && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-run-on-kubernetes" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-kubernetes into .cursor/skills/tao-run-on-kubernetes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-kubernetes", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-run-on-kubernetes--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-run-on-kubernetes --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-run-on-kubernetes .gemini/skills/tao-run-on-kubernetes && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-run-on-kubernetes" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-kubernetes into .gemini/skills/tao-run-on-kubernetes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-kubernetes", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-run-on-kubernetesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-run-on-kubernetes .github/skills/tao-run-on-kubernetes && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-run-on-kubernetes" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-kubernetes into .github/skills/tao-run-on-kubernetes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-kubernetes", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-run-on-kubernetes --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-run-on-kubernetes .opencode/skills/tao-run-on-kubernetes && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-run-on-kubernetes" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-kubernetes into .opencode/skills/tao-run-on-kubernetes/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-kubernetes", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-run-on-kubernetesKubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training.
Tao Run On Kubernetes is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training. Use when running on EKS / GKE / AKS / on-prem clusters with the NVIDIA GPU Operator installed, or when integrating TAO into an existing k8s-native ML platform.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including scripts and reference files (for example `BENCHMARK.md`, `config.json` and `config/skillspector-baseline.yaml`). Compatibility notes: Requires GPU worker nodes with NVIDIA driver branch 580, CUDA Toolkit 13.0, and NVIDIA Container Toolkit 1.19.0; a kubectl client authenticated to the…
It sits in DevOps & Cloud, covering Container orchestration. It works with Kubernetes, NVIDIA AI Platform and Google Kubernetes Engine. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
kubectlhelmbashshminikubeFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
docs.nvidia.comhelm.ngc.nvidia.comAlso links to:
kubernetes.iopytorch.orggithub.comkubeflow.orgvolcano.shkueue.sigs.k8s.ioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NGC_KEYACCESS_KEYSECRET_KEYHF_TOKENCRED_SECRETAWS_ACCESS_KEY_IDAWS_SECRET_ACCESS_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires GPU worker nodes with NVIDIA driver branch 580, CUDA Toolkit 13.0, and NVIDIA Container Toolkit 1.19.0; a `kubectl` client authenticated to the cluster; and the NVIDIA GPU Operator or device plugin. No nvidia-tao-sdk required — jobs are submitted with plain `kubectl`.
From compatibility in the SKILL.md frontmatter.
Tao Run On Kubernetes loads about 4.9k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 1,863 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
: no reachable cluster (kubeconfig at ~/.kube/config, \$KUBECONFIG, or in-pod service account)."- `~/.kube/config` — default discovery pathset -a; source /path/to/.env; set +a # omit if already exportedallowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 1,863 words, ~4,928 tokens.
.claude/skills/tao-run-on-kubernetes/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Submits TAO container jobs as Kubernetes Jobs. Works on any cluster reachable via kubeconfig (EKS / GKE / AKS / on-prem) or in-cluster service account (when running inside a pod).
Single-pod by default; opt into multi-node distributed training via num_nodes > 1 (uses Indexed Job + headless Service, see Multi-node training below).
Three checks: GPU host runtime ready, cluster reachable via kubectl, GPU
Operator/device plugin present.
# 0. GPU node host runtime.
# Run this on each self-managed GPU worker node or in the node image build.
# Set TAO_K8S_SKIP_NODE_RUNTIME_CHECK=1 only when using managed GPU nodes whose
# driver/toolkit lifecycle is owned by the cloud provider or GPU Operator policy.
if [ "${TAO_K8S_SKIP_NODE_RUNTIME_CHECK:-0}" != "1" ]; then
TAO_SKILL_BANK_ROOT="${TAO_SKILL_BANK_ROOT:-$PWD}"
SETUP_SCRIPT="${TAO_SKILL_BANK_ROOT}/skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh"
bash "$SETUP_SCRIPT" --backend kubernetes --check-only || {
echo "MISSING: TAO Kubernetes GPU node runtime is not ready."
echo "For self-managed GPU nodes, run after user approval:"
echo " bash \"$SETUP_SCRIPT\" --backend kubernetes --install --yes"
echo "For managed clusters, verify the node image/GPU Operator policy installs driver 580 and toolkit 1.19.0, then set TAO_K8S_SKIP_NODE_RUNTIME_CHECK=1."
exit 1
}
fi
# 1. Cluster reachable (kubeconfig OR in-cluster service account)
command -v kubectl >/dev/null 2>&1 || {
echo "MISSING: kubectl not found on PATH. Install kubectl to submit Jobs."
exit 1
}
kubectl cluster-info >/dev/null 2>&1 || {
echo "MISSING: no reachable cluster (kubeconfig at ~/.kube/config, \$KUBECONFIG, or in-pod service account)."
echo "Configure kubectl for your cluster, or set \$KUBECONFIG:"
echo " EKS: aws eks update-kubeconfig --name <cluster> --region <region>"
echo " GKE: gcloud container clusters get-credentials <cluster> --region <region>"
echo " AKS: az aks get-credentials --resource-group <rg> --name <cluster>"
echo " local: minikube start (see 'Local cluster' below)"
exit 1
}
# 2. NVIDIA GPU Operator present (soft check — warn, don't fail)
gpu=$(kubectl get nodes -o jsonpath='{range .items[*]}{.status.allocatable.nvidia\.com/gpu}{"\n"}{end}' 2>/dev/null | grep -v '^$' | head -1)
if [ -z "$gpu" ] || [ "$gpu" = "0" ]; then
echo "WARN: no nvidia.com/gpu allocatable on this cluster."
echo "Install the NVIDIA GPU Operator before submitting GPU jobs:"
echo " https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/getting-started.html"
fiThe GPU node runtime check is mandatory for self-managed nodes. For managed
clusters where the client is not running on a GPU worker, verify the provider
node image or GPU Operator policy and set TAO_K8S_SKIP_NODE_RUNTIME_CHECK=1
instead of running the installer on the client. The GPU-capacity warning here is
a soft check; the submit verb re-checks allocatable nvidia.com/gpu and
hard-fails before applying the manifest (there is no gang scheduling, so a
too-big Job would sit Pending forever).
~/.kube/config — default discovery path$KUBECONFIG — alternate pathdefault.imagePullSecrets.envFrom.secretRef), never inline. Legacy ACCESS_KEY/SECRET_KEY are mapped by tao-data-io.Do not ask for Brev or SLURM credentials for Kubernetes runs. Ask for
S3 credentials only when the selected workflow uses s3:// inputs or outputs,
and ask for model-specific credentials such as HF_TOKEN only when the selected
model requires them. Before launch, verify the selected namespace can create
Jobs, dataset/result paths are visible from the pod, and PVC/mounted filesystem
paths are proven to be mounted into the job container; an agent-host local path
is not sufficient proof.
tao-run-on-kubernetes is a platform consumer: it runs a spec-bundle via
kubectl, mutating only the job-record. No nvidia-tao-sdk, no tao_sdk import —
jobs are submitted with plain kubectl apply.
$BANK = ${TAO_SKILL_BANK_PATH}.
GPU-capacity gate — hard-fail first (no gang scheduling → a too-big Job
sits Pending forever):
ALLOC=$(kubectl get nodes -o jsonpath='{range .items[*]}{.status.allocatable.nvidia\.com/gpu}{"\n"}{end}' | awk '{s+=$1} END{print s+0}')
[ "${ALLOC:-0}" -ge "$NUM_GPUS" ] || { echo "insufficient GPUs: need $NUM_GPUS, allocatable $ALLOC"; exit 1; }Storage tier (via tao-data-io): A = mount a bound PVC/NFS holding the
data (author the mount paths, no fetch — the air-gap answer, and what the
packaged template does); C = ephemeral: an initContainer fetches from S3
into a shared emptyDir and a final step uploads results to S3 before TTL.
Tier C holds the GPU while it downloads. A pod reserves nvidia.com/gpu
for its whole lifetime, initContainers included, so a large tier-C fetch —
or a first-time multi-GB image pull — is billed and reaper-eligible idle GPU
time, exactly like pulling inside a SLURM allocation. Prefer tier A when the
data is already on a PVC; choose tier C knowingly, for small inputs.
A producer action request may declare several mounts, including duplicate-
source aliases for logical and embedded absolute paths. Read
references/action-request.md and use its staging map plus packaged
renderer. For mode=config, materialize the producer's nested spec first,
stage that exact generated file, and pass it back to the renderer so the
pod receives a verified read-only config mount. The legacy single-root
template cannot represent that contract.
Open the record — mints the id, binds results_dir, before launch:
JOB_ID=$("$BANK/scripts/tao_job_record.py" open --platform kubernetes --image "$IMAGE" \
--network-arch "$ARCH" --action "$ACTION" --storage-tier "$TIER" --results-dir "$RESULTS_DIR")results_dir must be a mounted (surviving) volume path or an S3 prefix —
ttlSecondsAfterFinished deletes the Job and its logs after it ends, so
nothing is recoverable from the Job object later.
Render, gate, apply, and record RUNNING. For a producer action request,
follow references/action-request.md; it owns backend-name normalization,
conditional Secret references, native argv rendering, server dry-run, and
binding the applied object name to the job-record. For a simple one-root
spec-bundle, render templates/k8s/single-pod-job.yaml.tmpl, run
redact_secrets.py lint plus kubectl apply --dry-run=server, apply it, and
mark the record with backend-ref=<namespace>/<actual-object-name>.
A submit that skipped the gate or the open has no id — so it cannot launch.
Keep K8S_JOB_NAME from submit. On reattach, read the job-record's
backend_ref=<namespace>/<name> and recover both values from that field; do
not assume the Kubernetes name equals the record id.
kubectl get job "$K8S_JOB_NAME" -n "$NAMESPACE" \
-o jsonpath='{.status.conditions[0].type} {.status.active} {.status.succeeded} {.status.failed}'| kubectl signal | vocab |
|---|---|
| no pods scheduled | PENDING (kubectl get pods -n "$NAMESPACE" -l job-name="$K8S_JOB_NAME" → ImagePullBackOff / Insufficient nvidia.com/gpu in message) |
active ≥ 1 | RUNNING |
condition Complete | COMPLETE |
condition Failed | ERROR (classify from the pod's terminated reason — OOMKilled → ERR_INFRA) |
| Job/pod not found | UNKNOWN (may be TTL-deleted — the job-record is the source of truth) |
kubectl logs -n "$NAMESPACE" -l "job-name=$K8S_JOB_NAME" --tail "${N:-200}"kubectl delete job "$K8S_JOB_NAME" -n "$NAMESPACE" --cascade=foreground
if [ -n "${CRED_SECRET:-}" ]; then
kubectl delete secret "$CRED_SECRET" -n "$NAMESPACE" --ignore-not-found
fi
"$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state CANCELED --source agentSame four verbs, plus:
kubectl version -o json) — the pod
hostname <job>-<index> (PodIndexLabel) that MASTER_ADDR=<job>-0.<svc>
resolves to needs it; on older clusters rank-0 hangs at rendezvous.gpus_per_node × nodes (no gang scheduling → a partial start leaves rank-0 waiting forever).templates/k8s/indexed-job.yaml.tmpl — the headless Service +
Indexed Job + rendezvous env (WORLD_SIZE = node count, NODE_RANK from
JOB_COMPLETION_INDEX, MASTER_ADDR=<job>-0.<svc>, /dev/shm 16Gi so NCCL
doesn't silently hang). kubectl apply -f creates the Service and Job together;
cancel deletes the Job (Foreground) and the Service.A throwaway minikube/kind cluster exercises admission, the four verbs,
job-record wiring, and log plumbing without cluster quota — and is what an
agent-driven eval should provision for itself. kubectl and minikube are
single static binaries needing no root, so a non-root CI container can install
them itself.
Two prerequisites keep a rendered Job Pending, and the first masks the second:
the PVC the template mounts must exist (persistentvolumeclaim "<name>" not found fires before any GPU complaint), then a Job requesting nvidia.com/gpu
on a GPU-less cluster reports Insufficient nvidia.com/gpu and waits forever.
Render NUM_GPUS=0 for a lifecycle-only run and say GPU scheduling was not
verified; on a Linux GPU host, minikube start --driver=docker --gpus all
passes real GPUs through, so one GPU box suffices for a GPU-real smoke.
Install commands, driver choice, the container/host-networking caveat, and the
fake-device-plugin middle option: references/local-cluster.md.
The simple single-pod template invokes its command via /bin/sh -c (POSIX sh,
present in busybox/distroless as well as TAO images). For producer action
requests, an args-mode command and its arguments are native container argv; a
producer that needs a shell declares the shell and its script explicitly. A
simple config-mode command also becomes native argv after {config_path} is
substituted. A producer-owned config command that is itself a multi-line shell
script is preserved verbatim under /bin/sh -c; the renderer never constructs
shell text from config values.
The submit verb refuses to launch GPU jobs on a cluster with no nvidia.com/gpu allocatable. For self-managed clusters, first run the tao-setup-nvidia-gpu-host install action on every GPU worker node or bake the same package set into the node image:
bash skills/platform/tao-setup-nvidia-gpu-host/scripts/setup-nvidia-gpu-host.sh --backend kubernetes --install --yesThen install the NVIDIA GPU Operator or device plugin:
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update
helm install --wait gpu-operator -n gpu-operator --create-namespace nvidia/gpu-operatorFull guide: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/getting-started.html
Set num_nodes > 1 (see the Multi-node (nodes > 1) verb
steps above) to run distributed training across N pods. Rendering
templates/k8s/indexed-job.yaml.tmpl provisions:
A headless Service named after the Job (selector: job-name=<job-name>, clusterIP: None, publishNotReadyAddresses: true so pods can rendezvous before they're all Ready).
An Indexed Job with parallelism = completions = num_nodes, completionMode: Indexed. Each pod gets JOB_COMPLETION_INDEX injected by k8s automatically (= the node rank).
A command wrapper that exports the rendezvous env vars before invoking the user command. Two naming conventions are exported simultaneously:
| Env var | Value | Read by |
|---|---|---|
WORLD_SIZE | num_nodes | TAO PyTorch container's nvidia_tao_pytorch/core/entrypoint.py (uses this to mean node count, even though PyTorch's own convention is total processes) |
NUM_GPU_PER_NODE | gpu_count | TAO PyTorch container's entrypoint |
NNODES | num_nodes | torchrun and PyTorch-standard rendezvous |
NPROC_PER_NODE | gpu_count | torchrun |
NODE_RANK | $JOB_COMPLETION_INDEX | both |
MASTER_ADDR | <job-name>-0.<job-name> (pod-0's DNS) | both |
MASTER_PORT | 29500 | both (TAO's default) |
Both naming conventions are set so TAO entrypoints (dino train, etc.) and raw torchrun commands work without modification.
For a TAO entrypoint, the container reads spec.train.num_nodes and the wired
env vars — e.g. dino train -e /tmp/spec.yaml with gpu_count=8, num_nodes=4
(4 × 8 = 32 GPUs total).
For raw torchrun-based commands (non-TAO containers), the wrapper invokes:
torchrun --nnodes=$NNODES --nproc-per-node=$NPROC_PER_NODE --node-rank=$NODE_RANK \
--master-addr=$MASTER_ADDR --master-port=$MASTER_PORT train.pyThe capacity check sums across nodes: gpu_count × num_nodes ≤ cluster's allocatable nvidia.com/gpu.
PodIndexLabel feature). On older clusters the MASTER_ADDR=<job>-0.<svc> DNS lookup fails. Verify with kubectl version.MASTER_PORT env var). Most CNIs (Calico, Cilium, AWS VPC CNI) allow this by default; restrictive NetworkPolicies must be relaxed.NCCL_SOCKET_IFNAME / NCCL_IB_HCA in the container env of the rendered manifest.For more sophisticated topologies (gang scheduling, PyTorch elastic / fault-tolerant training, MPI / Horovod, RDMA setup), reach for an operator instead of plain Indexed Job:
PyTorchJob, TFJob) — https://www.kubeflow.org/docs/components/training/ — for elastic PyTorch training with built-in restart logic.This skill's Indexed Job path is intentionally simple and dependency-free; if you need elastic restart or gang scheduling, layer one of these on top and submit jobs through the operator's CRD instead.
No nvidia.com/gpu resources allocatable on the cluster — the GPU Operator (or NVIDIA Device Plugin) isn't installed. Install per the link above; verify with kubectl get nodes -o jsonpath='{.items[*].status.allocatable}'.
ImagePullBackOff / ErrImagePull — the cluster can't pull the image. For nvcr.io: pre-create an image-pull secret in the namespace and reference it as the pod's imagePullSecrets in the rendered manifest:
Feed the key over stdin — --docker-password=$NGC_KEY would put the secret in
argv, where it is visible in the host's process table and shell history:
set -a; source /path/to/.env; set +a # omit if already exported
kubectl create secret generic ngc-pull-secret -n tao-jobs \
--type=kubernetes.io/dockerconfigjson \
--from-file=.dockerconfigjson=/dev/stdin <<EOF
{"auths": {"nvcr.io": {"username": "\$oauthtoken", "password": "${NGC_KEY}"}}}
EOF
# Verify without reading the secret back:
kubectl get secret ngc-pull-secret -n tao-jobs >/dev/null && echo SECRET_OKPod stays Pending forever — kubectl describe pod -l job-name=$JOB_ID shows the scheduling reason in the Events. Common causes: insufficient GPU capacity (Insufficient nvidia.com/gpu), no node matches the pod's nodeSelector, missing image-pull secret, or PVC mount failure.
OOMKilled (exit 137) — container exceeded memory. Reduce batch size, lower max_length, or add a memory request/limit and target a larger node.
CredentialError: Could not authenticate to a Kubernetes cluster — neither kubeconfig nor in-cluster auth worked. Run kubectl get nodes to verify your config, or set $KUBECONFIG to the right path.
backoff_limit=0 — failures fail the whole training run. For elastic restart (e.g., resume from checkpoint after a node death), use Kubeflow's PyTorchJob operator instead.MASTER_ADDR:MASTER_PORT).$NGC_KEY. You pre-create the secret in the target namespace and pass the name. K8s namespace conventions vary widely, so we keep secret creation explicit.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (scripts, references) in skills/tao-run-on-kubernetes of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Tao Run On Kubernetes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Run On Kubernetes this skillNVIDIA/skills | 3.6k | — | ~4.9k | Automated safety check: Warn | Apache-2.0 | |
| Aicr Analyzing SnapshotsNVIDIA/aicr | 440 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Mirrord Operatormetalbear-co/mirrord | 5.4k | 1 repos | ~4.6k | Automated safety check: Pass | MIT | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| KubeShark for KubernetesLukasNiessen/kubernetes-skill | 446 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Nim Operator InstallNVIDIA/k8s-nim-operator | 159 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 |
NVIDIA/aicr
A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…
metalbear-co/mirrord
Help users install and configure the mirrord Operator for team/enterprise environments.
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
LukasNiessen/kubernetes-skill
Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.
NVIDIA/k8s-nim-operator
Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional…
NVIDIA/k8s-nim-operator
Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and…
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Categories
Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training. Tao Run On Kubernetes is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training.
Tao Run On Kubernetes fits situations like: running on EKS / GKE / AKS / on-prem clusters with the NVIDIA GPU Operator installed; integrating TAO into an existing k8s-native ML platform.
Run `npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a claude-code`. Or copy the skill folder (skills/tao-run-on-kubernetes in NVIDIA/skills) into .claude/skills/tao-run-on-kubernetes in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a codex`. Or copy the skill folder (skills/tao-run-on-kubernetes in NVIDIA/skills) into .agents/skills/tao-run-on-kubernetes in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-run-on-kubernetes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-run-on-kubernetes, .gemini/skills/tao-run-on-kubernetes, .github/skills/tao-run-on-kubernetes and .opencode/skills/tao-run-on-kubernetes in your project.
Going by SKILL.md and its folder, Tao Run On Kubernetes needs Python for the scripts in its folder, the command-line tools its instructions call (kubectl, helm, bash, sh and minikube) and credentials named NGC_KEY, ACCESS_KEY, SECRET_KEY and HF_TOKEN. Our summary lists: Python 3; Docker; A credential in NGC_KEY; A credential in AWS_SECRET_ACCESS_KEY. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires GPU worker nodes with NVIDIA driver branch 580, CUDA Toolkit 13.0, and NVIDIA Container Toolkit 1.19.0; a `kubectl` client authenticated to the cluster; and the NVIDIA GPU Operator or device plugin. No nvidia-tao-sdk required — jobs are submitted with plain `kubectl`..
SKILL.md names 8 domains. In commands or code: docs.nvidia.com and helm.ngc.nvidia.com; the agent is likely to contact these when it follows the instructions. As links in the text: kubernetes.io, pytorch.org, github.com, kubeflow.org, volcano.sh and kueue.sigs.k8s.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 2 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Tao Run On Kubernetes is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Run On Kubernetes: Aicr Analyzing Snapshots (NVIDIA/aicr, 440 stars), Mirrord Operator (metalbear-co/mirrord, 5.4k stars), Devops (nicepkg/auto-company, 195 stars) and KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.