Kcli Cluster Deployment
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…
$ npx skills add NVIDIA/aicr --skill aicr-uat-report -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/aicr aicr-uat-report --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/aicr-uat-report .claude/skills/aicr-uat-report && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "aicr-uat-report" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-uat-report into .claude/skills/aicr-uat-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-uat-report", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-uat-reportType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/aicr --skill aicr-uat-report -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/aicr aicr-uat-report --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/aicr-uat-report .agents/skills/aicr-uat-report && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "aicr-uat-report" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-uat-report into .agents/skills/aicr-uat-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-uat-report", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/aicr --skill aicr-uat-report -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/aicr aicr-uat-report --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/aicr-uat-report .cursor/skills/aicr-uat-report && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "aicr-uat-report" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-uat-report into .cursor/skills/aicr-uat-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-uat-report", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/aicr.git --path .agents/skills/aicr-uat-report--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/aicr --skill aicr-uat-report -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/aicr aicr-uat-report --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/aicr-uat-report .gemini/skills/aicr-uat-report && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "aicr-uat-report" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-uat-report into .gemini/skills/aicr-uat-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-uat-report", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/aicr aicr-uat-reportInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/aicr --skill aicr-uat-report -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/aicr-uat-report .github/skills/aicr-uat-report && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "aicr-uat-report" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-uat-report into .github/skills/aicr-uat-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-uat-report", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/aicr --skill aicr-uat-report -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/aicr aicr-uat-report --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/aicr-uat-report .opencode/skills/aicr-uat-report && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "aicr-uat-report" agent skill from https://github.com/NVIDIA/aicr/tree/main/.agents/skills/aicr-uat-report into .opencode/skills/aicr-uat-report/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "aicr-uat-report", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
aicr-uat-reportA skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…
Aicr Uat Report is an agent skill from NVIDIA/aicr, published by the product's own GitHub organization. Use when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run workflow (uat-run.yaml). Triggers on "UAT report", "/aicr-uat-report", "which UAT combos are failing", "UAT pass rate", "download the UAT debug bundle", "why did the UAT run fail", or RC/release-candidate validation prep that needs the combinations to test manually. Runs the bundled uatreport.py, classifies failures as product vs infra…
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `uat_report.py`).
It sits in DevOps & Cloud, covering Container orchestration. It works with Google Kubernetes Engine, Amazon Web Services and Google Cloud. The repository describes itself as: Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e8f18da. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
ghpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Aicr Uat Report loads about 3.2k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 1,540 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/aicr at commit e8f18da, republished under its Apache-2.0 licence (© NVIDIA). 1,540 words, ~3,168 tokens.
.claude/skills/aicr-uat-report/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Reports how the UAT Run workflow
(https://github.com/NVIDIA/aicr/actions/workflows/uat-run.yaml) performed
over a lookback window, aggregated by service x GPU x intent, with each
failure classified as product signal (real test failure) or infra noise.
The output feeds the release process: combinations failing on main are
the ones to test more closely during RC validation.
/aicr-uat-report (optionally with a number of days)Do NOT use this skill to re-run, dispatch, or cancel UAT runs.
--days 7. Do not ask — default to 3 when
unspecified.main (empty aicr_version input) are reported by
default; that is the release-process signal. Add --all-versions only
if the user explicitly asks to compare against release-tag runs.--download-debug <dir> when the user
asks why something failed, or when Step 2 classifies a failure as product
signal. Skip it for a plain pass-rate report — bundles are tens of MB.The workflow's run-name encodes everything needed — no per-job digging:
UAT <reservation> <intent> @ <version|main>[ #dispatch_key].
Reservation names are <cloud>-<gpu> rows from
infra/uat/reservations.yaml (e.g. aws-h100, gcp-h100, azure-h100,
aws-gb200), and cloud maps to service: aws=EKS, gcp=GKE, azure=AKS,
kind=Kind (self-hosted nvkind lane). Runs before 2026-07-21 used a title
without the intent word; the script derives intent from the nightly
dispatch-key cell index (version-outer/intent-inner, training first) and
marks those rows "(intent derived)".
Each per-cloud workflow (uat-aws.yaml, uat-gcp.yaml, uat-azure.yaml,
uat-kind.yaml) runs tests/uat/<cloud>/run debug on failure — before
teardown, while the cluster is still up — and uploads
uat-<cloud>[-<intent>]-debug-<run_id> with 30-day retention. Reusable
workflows inherit the caller's run_id, so the artifact hangs off the
uat-run.yaml run the report already lists.
The upload is gated on failure() && steps.prep.outcome != 'skipped', so
no bundle exists when the cloud job died in bring-up or image build, or
when the failure was in a downstream job (evidence ingest) while the cluster
job passed. Absence is itself a classification signal, not an error.
What to open, in triage order (if-no-files-found: ignore, so any entry can
be missing; cluster-debug/ prefix omitted below):
| Open | When / what it answers |
|---|---|
MANIFEST.yaml | Always first — runId, config, resolved recipe + criteria, and failingChecks lifted from report.json |
report.json | Full validator results; absent if the run died before validate |
train-logs/**, serve-logs/** | A CUJ check failed (NCCL, inference-perf) |
readiness-gate.log | Readiness-gate failure — one ===== attempt N section per validate --phase deployment try; the last --- failed validator output (attempt N) --- block names each non-passed validator with its message and stdout, including the Failed resources: list. Runs before #2630 have no such block, only the raw validate output |
cr-skyhooks.yaml, node-reboot-fingerprint.txt | Tuning-race failure — Skyhook status.status, taints, bootID/kernel |
pods-notready.txt, events.txt | Scheduling, eviction, OOM |
logs-<namespace>.txt | The operator owning the failing resource |
nodes*, other cr-*.yaml, ns-*.txt | Broader node and operator state |
snapshot.yaml, recipe.yaml, dry-run.json | What was collected / resolved / deployed |
evidence-result.json, evidence/pointer.yaml | Signed-evidence emit outcome |
python3 .agents/skills/aicr-uat-report/uat_report.py --days 3It prints, per version, a Markdown table (Service | GPU | Intent | Pass | Failures) and a "Failure detail" section listing each failing run's
timestamp, URL, and the failed job/step names. It is read-only
(gh run list / gh run view). If gh is not authenticated, stop and
tell the user to run gh auth status.
Map the failed step name to a failure nature. This drives the RC priority ranking, so classify every failure:
| Failed step contains | Nature | Product signal? |
|---|---|---|
UAT - readiness gate | Deployed stack did not converge: a deployment-phase validator kept failing | YES — the likely owner is the component behind the failing validator(s), which the script prints as failing validators: |
UAT - validate, UAT - prep, CUJ/test phase names | Real test failure | YES — but the validate step also emits signed evidence, so confirm against report.json (Step 2b) before ranking |
Bringup Infra, provision/actuator steps | Infra bring-up failure | no |
Buildx, Build and push, image/GHCR steps | CI/image flake | no |
Validate inputs, UAT - install tagged [apply only: maybe infra] | helmfile apply or Argo CD sync failed | maybe — recurring = investigate |
UAT - install tagged [legacy apply+readiness: ambiguous …] | Pre-#2630 run: the step ran both the apply and the readiness gate | ambiguous — see below |
The script tags install and readiness failures in brackets. A job that has a
UAT - readiness gate step uses the split layout, so its install step is
apply-only. A job without one predates the split (#2630), and its
UAT - install (...) step — including kind's `UAT - install (helmfile apply
— covered both halves. For those legacy failures, pull the bundle (Step 2b): areadiness-gate.logwithresult=fail` attempts means the
gate failed (product signal); no gate log, or an empty one, means the apply
failed (maybe infra). Without a bundle, keep it ambiguous; do not guess.For readiness failures the failing validators: text comes from the
UAT readiness gate failed annotation the phase emits. When it is absent (the
annotation failed to post, or the run predates it), read the names from the
last failed-validator block in readiness-gate.log.
A retry that went green the same night (same combo, later timestamp, success) downgrades the earlier failure to a flake.
Skip this step entirely for a routine pass-rate report. Run it when the user asks why something failed, or when Step 2 found a test-phase failure worth root-causing:
python3 .agents/skills/aicr-uat-report/uat_report.py --days 3 \
--download-debug /tmp/uat-debug --max-downloads 3Prefer --run <id> (repeatable) over raising --max-downloads: usually only
the latest failure per failing combo is worth reading.
Bundles land in <dir>/<service>-<gpu>-<intent>-<run_id>/, each with a
printed digest — MANIFEST head, failing checks from report.json, and a
one-line contents summary. Read that summary for presence, not file names:
a missing evidence/ or report.json says the run died before that stage,
which is often the whole diagnosis. For a readiness-gate failure the digest
also echoes the last --- failed validator output block from
readiness-gate.log; its Failed resources: lines name the resources that
never became ready, so start with the logs of the operator that owns them.
Then open files per the Debug bundles table above.
Cite file_path:line and the run ID for every finding. Leave the download
directory in place — the user may want to keep digging.
Produce exactly two artifacts, in this order (see Output Format Reference). Sort the table worst-first: lowest pass ratio at the top; bold the Service/GPU/Intent cells of rows with product-signal failures. Include run URLs as links for at least the most recent failure of each failing combo.
A numbered priority list derived from the table:
nightly-intents: [],
kind lane) — expected churn, call out separately.## UAT report against `main` (<start>–<end>, N runs)
| Service | GPU | Intent | Pass | Failure nature |
|---|---|---|---|---|
| **AKS** | **H100** | **training** | **0/4** | Real test failures — every run fails at "UAT - validate (all phases)" ([latest](<url>)) |
| EKS | H100 | training | 2/5 | Infra only — 2x bring-up, 1x Buildx CI flake; no test-phase failures |
| GKE | H100 | training | 4/4 | Green |
## RC validation input
1. **<Service>/<GPU>/<intent> is the clear red flag** — <pass ratio,
failure signature, latest run link, whether it also fails on release
tags (env issue) or only main (regression candidate)>.
2. ...
N. **<green combos> are solidly green** — lowest manual-testing priority.Keep failure-nature cells to one sentence; detail beyond that belongs in the RC list, not the table.
gh run list returns nothing — window may predate retention or the
workflow was renamed; say so rather than reporting "all green".run-name format in
uat-run.yaml changed; read the workflow's current format string and
update NEW_TITLE/OLD_TITLE in uat_report.py in the same PR.infra/uat/reservations.yaml.cluster-debug/ missing or thin — the collector is best-effort and
its cloud credentials can expire on a long failure. Say the bundle is
incomplete; do not read it as the cluster being healthy.report.json — trust the bundle.
UAT - validate (all phases) + emit signed evidence covers two concerns,
so N/N passing checks under a failed step means the evidence leg failed,
not the product. Reclassify to infra before ranking it in Step 4.gh run view --log); it reports failed
step names and, on request, the uploaded cluster debug bundles--download-debug directory© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in .agents/skills/aicr-uat-report of NVIDIA/aicr.
Open the folder on GitHubat commit e8f18da
Aicr Uat Report next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Aicr Uat Report this skillNVIDIA/aicr | 440 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| Kcli Cluster Deploymentkarmab/kcli | 653 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| Apex Azure Cloud Migratejonathan-vella/apex | 217 | — | ~1.2k | Automated safety check: Pass | MIT | |
| Devopsnicepkg/auto-company | 195 | 2 repos | ~814 | Automated safety check: Pass | MIT | |
| Provider Bug Reviewmondoohq/mql | 412 | — | ~2.9k | Automated safety check: Pass | Custom licence | |
| Logfire Infrastructurepydantic/skills | 140 | — | ~1.8k | Automated safety check: Pass | MIT |
karmab/kcli
Guides deployment and management of Kubernetes clusters with kcli.
jonathan-vella/apex
WORKFLOW SKILL — Assess and migrate cross-cloud workloads to Azure: assessments and code conversion from AWS, GCP, Heroku, Kubernetes or Spring.
nicepkg/auto-company
Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).
mondoohq/mql
Deep static code review of an mql provider for logic errors, nil-handling bugs, pagination truncation, caching/id collisions, and other defects that silently give users wrong data.
pydantic/skills
Monitor hosts, Docker containers, Kubernetes clusters, database/queue/cache servers, and cloud-provider metrics with Pydantic Logfire — no application code required.
runwhen-contrib/runwhen-local
Add or enrich a resource type in an existing RunWhen Local discovery indexer (Azure azureapi, GCP gcpapi, AWS, or Kubernetes).
NVIDIA/aicr
Multi-agent PR review using Claude Code, Codex, and CodeRabbit.
NVIDIA/aicr
A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…
NVIDIA/aicr
Scaffolds an interactive guided demo script (demos/.sh), live or self-paced, with the Frame → Tell → Show → Close pattern.
NVIDIA/aicr
A skill your agent uses when building a self-contained HTML slide deck or visual talking-point for a technical concept or workflow (e.g.
NVIDIA/aicr
A skill your agent uses when drafting the human-readable GitHub release notes summary for an upcoming AICR release.
NVIDIA/aicr
A skill your agent uses when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which…
Categories
A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…. Aicr Uat Report is an agent skill from NVIDIA/aicr, published by the product's own GitHub organization.yaml).
Aicr Uat Report fits situations like: reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB20; X intent combinations are passing; failing in the UAT Run workflow (uat-run.yaml); /aicr-uat-report.
Run `npx skills add NVIDIA/aicr --skill aicr-uat-report -a claude-code`. Or copy the skill folder (.agents/skills/aicr-uat-report in NVIDIA/aicr) into .claude/skills/aicr-uat-report in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/aicr --skill aicr-uat-report -a codex`. Or copy the skill folder (.agents/skills/aicr-uat-report in NVIDIA/aicr) into .agents/skills/aicr-uat-report in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/aicr --skill aicr-uat-report -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aicr-uat-report, .gemini/skills/aicr-uat-report, .github/skills/aicr-uat-report and .opencode/skills/aicr-uat-report in your project.
Going by SKILL.md and its folder, Aicr Uat Report needs Python for the scripts in its folder and the command-line tools its instructions call (gh and python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Aicr Uat Report is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Aicr Uat Report: Kcli Cluster Deployment (karmab/kcli, 653 stars), Apex Azure Cloud Migrate (jonathan-vella/apex, 217 stars), Devops (nicepkg/auto-company, 195 stars) and Provider Bug Review (mondoohq/mql, 412 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/aicr, which has 440 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 10, 2026.
Source: NVIDIA/aicr on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.