Tao Run On Kubernetes
NVIDIA/skills
Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training.
A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…
$ npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/aicr aicr-analyzing-snapshots --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/aicr-analyzing-snapshots .claude/skills/aicr-analyzing-snapshots && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "aicr-analyzing-snapshots" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-analyzing-snapshots into .claude/skills/aicr-analyzing-snapshots/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-analyzing-snapshots", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-analyzing-snapshotsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/aicr aicr-analyzing-snapshots --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/aicr-analyzing-snapshots .agents/skills/aicr-analyzing-snapshots && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "aicr-analyzing-snapshots" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-analyzing-snapshots into .agents/skills/aicr-analyzing-snapshots/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-analyzing-snapshots", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/aicr aicr-analyzing-snapshots --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/aicr-analyzing-snapshots .cursor/skills/aicr-analyzing-snapshots && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "aicr-analyzing-snapshots" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-analyzing-snapshots into .cursor/skills/aicr-analyzing-snapshots/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-analyzing-snapshots", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/aicr.git --path .agents/skills/aicr-analyzing-snapshots--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/aicr aicr-analyzing-snapshots --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/aicr-analyzing-snapshots .gemini/skills/aicr-analyzing-snapshots && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "aicr-analyzing-snapshots" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-analyzing-snapshots into .gemini/skills/aicr-analyzing-snapshots/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-analyzing-snapshots", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/aicr aicr-analyzing-snapshotsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/aicr-analyzing-snapshots .github/skills/aicr-analyzing-snapshots && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "aicr-analyzing-snapshots" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-analyzing-snapshots into .github/skills/aicr-analyzing-snapshots/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-analyzing-snapshots", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/aicr aicr-analyzing-snapshots --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/aicr-analyzing-snapshots .opencode/skills/aicr-analyzing-snapshots && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "aicr-analyzing-snapshots" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-analyzing-snapshots into .opencode/skills/aicr-analyzing-snapshots/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-analyzing-snapshots", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
aicr-analyzing-snapshotsA skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…
Aicr Analyzing Snapshots is an agent skill from NVIDIA/aicr, published by the product's own GitHub organization. Use when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster assessment report from a snapshot. Triggers on: snapshot analysis, cluster review, provider comparison, GPU topology, node health, snapshot report.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in DevOps & Cloud, covering Container orchestration. It works with Kubernetes, NVIDIA AI Platform and Google Kubernetes Engine. The repository describes itself as: Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 633c358. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python and bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Aicr Analyzing Snapshots loads about 3.5k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,254 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/aicr at commit 633c358, republished under its Apache-2.0 licence (© NVIDIA). 1,254 words, ~3,473 tokens.
.claude/skills/aicr-analyzing-snapshots/SKILL.md (or your agent's skills folder).Systematic analysis of AICR snapshot YAML files to extract cluster identity, provider characteristics, GPU topology, node health, software stack, and operational signals. Produces a structured Markdown report.
Snapshot files are large (50K-80K+ tokens). Never read the whole file.
Use mcp__plugin_context-mode_context-mode__execute_file with Python/YAML
parsing to extract sections, or use targeted Read with offset/limit on
specific line ranges found via Grep.
import yaml
data = yaml.safe_load(FILE_CONTENT)
meta = data.get('metadata', {})
measurements = data.get('measurements', [])
print("=== METADATA ===")
for k, v in meta.items():
print(f" {k}: {v}")
print("\n=== MEASUREMENTS ===")
for m in measurements:
subtypes = [s.get('subtype', s.get('name', '?')) for s in m.get('subtypes', [])]
print(f" {m['type']}: {subtypes}")Key fields for provider identification:
| Field Path | What It Reveals |
|---|---|
K8s.server.version | K8s version + vendor suffix (-eks-, -gke, -aks, +lke) |
K8s.node.provider | Mapped provider: eks, gke, aks, oke, lke, metal3, kind |
K8s.node.provider-id | Raw provider URI (aws://, gce://, azure://, oci://, linode://, metal3://) |
K8s.node.kernel-version | Kernel + arch indicator (e.g., -64k = ARM 64K pages) |
K8s.node.container-runtime-* | Runtime name and version |
K8s.node.kubelet-version | Kubelet version |
K8s.node.os-image | OS description string |
Provider detection logic:
| provider-id prefix | Service | Notes |
|---|---|---|
aws:// | eks | Amazon EKS |
gce:// | gke | Google GKE |
azure:// | aks | Azure AKS |
oci:// | oke | Oracle OKE |
linode:// | lke | Akamai Cloud / Linode LKE |
metal3:// | bare-metal | Metal3/Ironic, self-managed |
kind:// | kind | Local dev cluster |
| (none/other) | any | Self-managed, check version string |
If provider-id is absent, check K8s.server.version for vendor substrings.
Key fields from GPU.smi:
| Field | Example | Significance |
|---|---|---|
gpu.model | NVIDIA GB300 | Maps to accelerator criteria |
gpu.product-architecture | Blackwell | GPU generation |
gpu-count | 4 | GPUs per node |
driver | 580.126.16 | NVIDIA driver version |
cuda-version | 13.0 | CUDA toolkit version |
gpu.addressing-mode | ATS | ATS = unified CPU-GPU memory (Grace) |
gpu.persistence-mode | Disabled/Enabled | Should be Enabled for production |
gpu.vbios-version | 97.10.4A.00.1A | Firmware version |
gpu.gsp-firmware-version | 580.126.16 | GSP firmware |
Accelerator mapping (checked in order, case-insensitive):
| gpu.model contains | Accelerator |
|---|---|
gb200 | gb200 (check before b200) |
gb300 | gb200 class (Blackwell NVL family) |
b200 | b200 |
h100 | h100 |
gh200 | unresolved — Grace Hopper Superchip, not the discrete H200 GPU (check before h200) |
h200 | h200 (discrete H200 GPU) |
a100 | a100 |
l40s | l40s |
l40 | l40 |
rtx pro 6000 | rtx-pro-6000 |
From OS.release: ID, VERSION_ID, PRETTY_NAME
From OS.grub: Boot parameters (check for iommu, console, init_on_free)
From OS.kmod: Loaded kernel modules (look for nvidia*, nv_peer_mem,
gdrdrv, ib_*, mlx5_* for RDMA/InfiniBand)
From OS.sysctl (key tuning parameters):
| Sysctl | Good Value for GPU | Why |
|---|---|---|
vm.swappiness | <= 10 | Minimize swapping for GPU workloads |
vm.overcommit_memory | 1 | Allow overcommit for training |
vm.nr_hugepages | > 0 (ideal) | Large page performance |
fs.file-max | High (9223372036854775807) | Sufficient file descriptors |
kernel.threads-max | > 1M | Sufficient threads |
vm.min_free_kbytes | > 1M | Memory reserve |
From NodeTopology.summary: node-count, taint-count, label-count
From NodeTopology.taint and NodeTopology.label, read the items list — one
entry per distinct reading, sorted by key/value (taints: key/effect/value):
| Item Field | What It Holds |
|---|---|
context.key | Taint or label key, verbatim |
context.value | Taint or label value (may be empty) |
context.effect | Taints only: NoSchedule, PreferNoSchedule, NoExecute |
data.node-count | True node total, including nodes dropped by truncation |
data.node-list | Comma-separated node names (one of node-list / node-list-ref) |
data.node-list-ref | Key into the subtype's data map whose entry holds the names (one of node-list / node-list-ref) |
data.truncated | true when the node list is capped and ends with (+N more) |
Current snapshots also carry the older data map on both subtypes; items is
authoritative. Take counts from data.node-count rather than splitting
node-list, and read data.truncated rather than probing for a (+N more)
suffix. Use topology.LabelReadings / TaintReadings to resolve items into
hydrated readings — they expand node-list-ref automatically, so callers do
not need to implement the reference logic themselves.
Older snapshots (no items): fall back to the folded data map —
effect|value|node1,node2,... for taints, value|node1,node2,... for labels.
That encoding is lossy, so qualify anything derived from it:
<key>.<value>, indistinguishable from a label
literally named that, and one of the colliding readings is dropped. Report
such a key verbatim instead of asserting a key/value split..<effect> and its value has
only two fields (value|nodes); two taints sharing key and effect collapse
into one entry.summary.taint-count / label-count count map entries there, so they
under-report wherever a collapse occurred, and node counts reflect only what
survived truncation.High-value labels to extract (skip feature.node.kubernetes.io/cpu-cpuid.*):
| Label Prefix | What It Reveals |
|---|---|
kubernetes.io/arch.* | CPU architecture (amd64 vs arm64 = heterogeneous) |
nvidia.com/gpu.* | GPU product, family, memory, compute, count, MIG state |
nvidia.com/cuda.* | CUDA driver/runtime versions |
nvidia.com/mig.* | MIG capable/config/strategy |
nvidia.com/gpu.clique.* | NVLink GPU cliques (multi-node NVLink domains) |
resource.nvidia.com/computeDomain | Unified compute domain |
network.topology.nvidia.com/accelerator.* | NVLink fabric blocks |
node-type.* | Hardware type (gb300, standard) |
node-pool.* | Pool assignment (gpu-pool, cpu-pool) |
node.dgxc.nvidia.com/* | DGX Cloud node classification |
k8saas.nvidia.com/* | K8SaaS management (NVSentinel cordon/uncordon) |
dgxc.nvidia.com/nvsentinel-state | Health state (remediation-failed, healthy) |
nvsentinel.dgxc.nvidia.com/* | NVSentinel component versions, driver state |
network.nvidia.com/operator.* | Network operator MOFED/NIC config state |
metal3.io/uuid.* | Metal3 bare-metal node UUIDs |
workload.* | Workload type (gpu, general) |
feature.node.kubernetes.io/rdma.* | RDMA available/capable |
feature.node.kubernetes.io/network-sriov.* | SR-IOV capability |
feature.node.kubernetes.io/pci-15b3.* | Mellanox ConnectX presence |
feature.node.kubernetes.io/pci-10de.* | NVIDIA GPU PCI presence |
nvidia.com/dra-kubelet-plugin | DRA (Dynamic Resource Allocation) |
From K8s.image: All deployed container images and versions.
From K8s.policy: Flattened GPU Operator ClusterPolicy spec (dot-notation).
Key policy fields:
| Policy Field | What to Check |
|---|---|
driver.enabled | GPU driver managed by operator |
driver.version | Driver version in policy |
driver.rdma.enabled | RDMA support |
toolkit.enabled | Container toolkit |
devicePlugin.enabled | Device plugin active |
dcgm.enabled / dcgmExporter.enabled | GPU monitoring |
migManager.enabled | MIG management |
ccManager.enabled / ccManager.defaultMode | Confidential Computing |
sandboxWorkloads.enabled | Sandbox/KubeVirt workloads |
psa.enabled | Pod Security Admission |
vfioManager.enabled | VFIO passthrough |
From K8s.slinky-slurm, report:
collection-state: absent, detected, unsupported-multicluster, or
unknowndetected means a Controller declaration exists, not that Slurm or its
operator is healthy. Child items and counts are emitted only after all required
APIs and references are collected conclusively; their absence is otherwise not
confirmed absence. Never infer platform: slurm from this subtype.
From K8s.mariadb-operator, report collection-state as official
MariaDB-operator API conflict evidence:
absent: official API group conclusively absentapi-detected: official API footprint present without observed MariaDB CRscrs-detected: one or more official MariaDB CRs observedunknown: discovery or List was inconclusiveThese states do not prove database availability, operator health, or the
existence of an external database such as RDS. Never infer
accounting.databaseSource.
From SystemD.containerd.service, SystemD.kubelet.service, SystemD.docker.service:
| Field | What to Check |
|---|---|
ActiveState | Should be active |
SubState | Should be running |
LimitNOFILE | File descriptor limits |
LimitMEMLOCK | Memory lock limits (important for RDMA) |
KillMode | process for containerd (graceful) |
Delegate | true for containerd (cgroup delegation) |
CPUAccounting | Resource accounting |
Structure the output as:
# Snapshot Analysis: {name}
> Source: {file} | Captured: {timestamp} | AICR: {version}
## Cluster Identity
Table: source-node, provider, K8s version, node count, GPU model, total GPUs
## Provider-Differentiating Insights
### 1. Provider Type (cloud vs bare-metal, managed vs self-managed)
### 2. CPU Architecture (homogeneous vs heterogeneous, ARM vs x86)
### 3. GPU Hardware (model, architecture, memory, driver, CUDA, MIG, persistence)
### 4. Network Topology (NVLink blocks, cliques, compute domains, RDMA, SR-IOV)
### 5. Management Layer (K8SaaS, NVSentinel health, cordon state)
### 6. Job Scheduling (Slurm/Slinky presence, HPC vs cloud-native)
### 7. Networking Stack (CNI, RDMA, SR-IOV, DOCA/MOFED)
### 8. Security (Confidential Computing, PSA, DRA)
### 9. Operational Signals (sysctl tuning, hugepages, persistence mode)
## Software Stack
### Key Container Images (table)
### OS and Kernel (table)
## Node Inventory
List nodes by rack/block/pool
## Operational Flags
Anything unusual: GPU health issues, disabled persistence mode,
missing hugepages, NVSentinel remediation failures, etc.metal3:// provider-id with per-node UUIDsk8saas.nvidia.com/* management labelsAfter analysis, map the snapshot to AICR recipe criteria:
aicr recipe \
--service {detected_service} \
--accelerator {detected_accelerator} \
--os {detected_os} \
--intent {training|inference} \
--snapshot {snapshot_file}| Criteria | Extracted From | Valid Values |
|---|---|---|
| service | K8s.node.provider / K8s.server.version | eks, gke, aks, oke, kind, lke |
| accelerator | GPU.smi.gpu.model | h100, h200, gb200, b200, a100, l40s, l40, rtx-pro-6000 |
| os | OS.release.ID | ubuntu, rhel, cos, amazonlinux, talos, ol |
| intent | User-specified | training, inference |
| platform | User-specified | dynamo, kubeflow, nim, runai, slurm |
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/aicr-analyzing-snapshots of NVIDIA/aicr.
Open the folder on GitHubat commit 633c358
Aicr Analyzing Snapshots next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Aicr Analyzing Snapshots this skillNVIDIA/aicr | 440 | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Tao Run On KubernetesNVIDIA/skills | 3.5k | — | ~4.9k | Automated safety check: Warn | Apache-2.0 | |
| Devopsnicepkg/auto-company | 194 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| KubeShark for KubernetesLukasNiessen/kubernetes-skill | 446 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Nim Operator InstallNVIDIA/k8s-nim-operator | 159 | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| Nim Operator UninstallNVIDIA/k8s-nim-operator | 159 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 |
NVIDIA/skills
Kubernetes execution platform — submits TAO container jobs as k8s Jobs with NVIDIA GPU scheduling; single-pod for one node, Indexed Jobs for multi-node distributed training.
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
LukasNiessen/kubernetes-skill
Keeps Kubernetes manifests, Helm charts and policies grounded by diagnosing six failure modes, such as insecure defaults and API drift, and loading only matching references.
NVIDIA/k8s-nim-operator
Install NVIDIA NIM Operator on Kubernetes with prerequisite checks, optional NVIDIA GPU Operator dependency installation, public or local Helm chart selection, optional Dynamo support, and optional…
NVIDIA/k8s-nim-operator
Safely uninstall NVIDIA NIM Operator from Kubernetes with inventory checks, explicit approval gates for destructive actions, optional custom resource cleanup, optional CRD removal, and…
NVIDIA/dcgm-exporter
A skill your agent uses when validating DCGM Exporter in a local GPU-backed k3d/Kubernetes environment.
NVIDIA/aicr
Multi-agent PR review using Claude Code, Codex, and CodeRabbit.
NVIDIA/aicr
Scaffolds an interactive guided demo script (demos/.sh), live or self-paced, with the Frame → Tell → Show → Close pattern.
NVIDIA/aicr
A skill your agent uses when building a self-contained HTML slide deck or visual talking-point for a technical concept or workflow (e.g.
NVIDIA/aicr
A skill your agent uses when drafting the human-readable GitHub release notes summary for an upcoming AICR release.
NVIDIA/aicr
A skill your agent uses when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which…
NVIDIA/aicr
A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…
Categories
A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…. Aicr Analyzing Snapshots is an agent skill from NVIDIA/aicr, published by the product's own GitHub organization. Use when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster assessment report from a snapshot.
Aicr Analyzing Snapshots fits situations like: analyzing an AICR snapshot YAML file; reviewing cluster state; comparing provider characteristics; extracting GPU/network topology insights.
Run `npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a claude-code`. Or copy the skill folder (.agents/skills/aicr-analyzing-snapshots in NVIDIA/aicr) into .claude/skills/aicr-analyzing-snapshots in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a codex`. Or copy the skill folder (.agents/skills/aicr-analyzing-snapshots in NVIDIA/aicr) into .agents/skills/aicr-analyzing-snapshots in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/aicr --skill aicr-analyzing-snapshots -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aicr-analyzing-snapshots, .gemini/skills/aicr-analyzing-snapshots, .github/skills/aicr-analyzing-snapshots and .opencode/skills/aicr-analyzing-snapshots in your project.
SKILL.md names no scripts, command-line tools or credentials: Aicr Analyzing Snapshots is instructions for the agent only. Our summary lists: Python 3; Docker.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Aicr Analyzing Snapshots is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Aicr Analyzing Snapshots: Tao Run On Kubernetes (NVIDIA/skills, 3.5k stars), Devops (nicepkg/auto-company, 194 stars), KubeShark for Kubernetes (LukasNiessen/kubernetes-skill, 446 stars) and Nim Operator Install (NVIDIA/k8s-nim-operator, 159 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/aicr, which has 440 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/aicr on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.