Official agent skill

Aicr Uat Report

by NVIDIA in NVIDIA/aicr

A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…

OfficialApache-2.0Auto-check passedDevOps & Cloud

Install Aicr Uat Report

skills CLI
$ npx skills add NVIDIA/aicr --skill aicr-uat-report -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/aicr aicr-uat-report --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/aicr.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/aicr-uat-report .claude/skills/aicr-uat-report && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
aicr-uat-report
GitHub stars
440
Token cost
~3.2k tokens
SKILL.md length
1,540 words
Files
2
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…

  • Works in 4 steps: Run the report script → Classify each failure → Render the report → …
  • Reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB20
  • SKILL.md covers When to Use, Inputs, How the Data Works and Procedure, plus 3 more sections
  • Runs Python scripts from its folder; calls gh and python3

What it does

Aicr Uat Report is an agent skill from NVIDIA/aicr, published by the product's own GitHub organization. Use when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run workflow (uat-run.yaml). Triggers on "UAT report", "/aicr-uat-report", "which UAT combos are failing", "UAT pass rate", "download the UAT debug bundle", "why did the UAT run fail", or RC/release-candidate validation prep that needs the combinations to test manually. Runs the bundled uatreport.py, classifies failures as product vs infra…

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `uat_report.py`).

It sits in DevOps & Cloud, covering Container orchestration. It works with Google Kubernetes Engine, Amazon Web Services and Google Cloud. The repository describes itself as: Tooling for optimized, validated, and reproducible GPU-accelerated AI runtime in Kubernetes. The licence is Apache-2.0.

When your agent uses it

  • Reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB20
  • X intent combinations are passing
  • Failing in the UAT Run workflow (uat-run.yaml)
  • /aicr-uat-report

Example prompts

  • “UAT report”
  • “/aicr-uat-report”
  • “which UAT combos are failing”
  • “/aicr-uat-report”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Run the report script
  2. Classify each failure
  3. Render the report
  4. Write the RC validation input

What it can do on your machine

Read from SKILL.md and the folder at commit e8f18da. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • gh
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Aicr Uat Report loads about 3.2k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 1,540 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/aicr at commit e8f18da, republished under its Apache-2.0 licence (© NVIDIA). 1,540 words, ~3,168 tokens.

Download SKILL.mdSave it as .claude/skills/aicr-uat-report/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
aicr-uat-report
description
Use when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run workflow (uat-run.yaml). Triggers on "UAT report", "/aicr-uat-report", "which UAT combos are failing", "UAT pass rate", "download the UAT debug bundle", "why did the UAT run fail", or RC/release-candidate validation prep that needs the combinations to test manually. Runs the bundled uat_report.py, classifies failures as product vs infra signal, prints a summary table plus an RC validation priority list, and can download the per-run cluster debug bundles for triage.

AICR UAT Report

Reports how the UAT Run workflow (https://github.com/NVIDIA/aicr/actions/workflows/uat-run.yaml) performed over a lookback window, aggregated by service x GPU x intent, with each failure classified as product signal (real test failure) or infra noise. The output feeds the release process: combinations failing on main are the ones to test more closely during RC validation.

When to Use

  • User asks for a UAT report, UAT pass rates, or failing UAT combinations
  • User invokes /aicr-uat-report (optionally with a number of days)
  • Release prep needs the list of service/GPU combos to validate manually
  • User asks why a UAT run failed, or for the debug bundle behind a failure

Do NOT use this skill to re-run, dispatch, or cancel UAT runs.

Inputs

  • days (optional, default 3): lookback window. If the user says "over the last week", pass --days 7. Do not ask — default to 3 when unspecified.
  • Only runs against main (empty aicr_version input) are reported by default; that is the release-process signal. Add --all-versions only if the user explicitly asks to compare against release-tag runs.
  • debug bundles (optional): pass --download-debug <dir> when the user asks why something failed, or when Step 2 classifies a failure as product signal. Skip it for a plain pass-rate report — bundles are tens of MB.

How the Data Works

The workflow's run-name encodes everything needed — no per-job digging: UAT <reservation> <intent> @ <version|main>[ #dispatch_key]. Reservation names are <cloud>-<gpu> rows from infra/uat/reservations.yaml (e.g. aws-h100, gcp-h100, azure-h100, aws-gb200), and cloud maps to service: aws=EKS, gcp=GKE, azure=AKS, kind=Kind (self-hosted nvkind lane). Runs before 2026-07-21 used a title without the intent word; the script derives intent from the nightly dispatch-key cell index (version-outer/intent-inner, training first) and marks those rows "(intent derived)".

Debug bundles

Each per-cloud workflow (uat-aws.yaml, uat-gcp.yaml, uat-azure.yaml, uat-kind.yaml) runs tests/uat/<cloud>/run debug on failure — before teardown, while the cluster is still up — and uploads uat-<cloud>[-<intent>]-debug-<run_id> with 30-day retention. Reusable workflows inherit the caller's run_id, so the artifact hangs off the uat-run.yaml run the report already lists.

The upload is gated on failure() && steps.prep.outcome != 'skipped', so no bundle exists when the cloud job died in bring-up or image build, or when the failure was in a downstream job (evidence ingest) while the cluster job passed. Absence is itself a classification signal, not an error.

What to open, in triage order (if-no-files-found: ignore, so any entry can be missing; cluster-debug/ prefix omitted below):

OpenWhen / what it answers
MANIFEST.yamlAlways first — runId, config, resolved recipe + criteria, and failingChecks lifted from report.json
report.jsonFull validator results; absent if the run died before validate
train-logs/**, serve-logs/**A CUJ check failed (NCCL, inference-perf)
readiness-gate.logReadiness-gate failure — one ===== attempt N section per validate --phase deployment try; the last --- failed validator output (attempt N) --- block names each non-passed validator with its message and stdout, including the Failed resources: list. Runs before #2630 have no such block, only the raw validate output
cr-skyhooks.yaml, node-reboot-fingerprint.txtTuning-race failure — Skyhook status.status, taints, bootID/kernel
pods-notready.txt, events.txtScheduling, eviction, OOM
logs-<namespace>.txtThe operator owning the failing resource
nodes*, other cr-*.yaml, ns-*.txtBroader node and operator state
snapshot.yaml, recipe.yaml, dry-run.jsonWhat was collected / resolved / deployed
evidence-result.json, evidence/pointer.yamlSigned-evidence emit outcome

Procedure

Step 1 — Run the report script
bash
python3 .agents/skills/aicr-uat-report/uat_report.py --days 3

It prints, per version, a Markdown table (Service | GPU | Intent | Pass | Failures) and a "Failure detail" section listing each failing run's timestamp, URL, and the failed job/step names. It is read-only (gh run list / gh run view). If gh is not authenticated, stop and tell the user to run gh auth status.

Step 2 — Classify each failure

Map the failed step name to a failure nature. This drives the RC priority ranking, so classify every failure:

Failed step containsNatureProduct signal?
UAT - readiness gateDeployed stack did not converge: a deployment-phase validator kept failingYES — the likely owner is the component behind the failing validator(s), which the script prints as failing validators:
UAT - validate, UAT - prep, CUJ/test phase namesReal test failureYES — but the validate step also emits signed evidence, so confirm against report.json (Step 2b) before ranking
Bringup Infra, provision/actuator stepsInfra bring-up failureno
Buildx, Build and push, image/GHCR stepsCI/image flakeno
Validate inputs, UAT - install tagged [apply only: maybe infra]helmfile apply or Argo CD sync failedmaybe — recurring = investigate
UAT - install tagged [legacy apply+readiness: ambiguous …]Pre-#2630 run: the step ran both the apply and the readiness gateambiguous — see below

The script tags install and readiness failures in brackets. A job that has a UAT - readiness gate step uses the split layout, so its install step is apply-only. A job without one predates the split (#2630), and its UAT - install (...) step — including kind's `UAT - install (helmfile apply

  • readiness gate)— covered both halves. For those legacy failures, pull the bundle (Step 2b): areadiness-gate.logwithresult=fail` attempts means the gate failed (product signal); no gate log, or an empty one, means the apply failed (maybe infra). Without a bundle, keep it ambiguous; do not guess.

For readiness failures the failing validators: text comes from the UAT readiness gate failed annotation the phase emits. When it is absent (the annotation failed to post, or the run predates it), read the names from the last failed-validator block in readiness-gate.log.

A retry that went green the same night (same combo, later timestamp, success) downgrades the earlier failure to a flake.

Show full SKILL.md (665 more words)Show less
Step 2b — Pull debug bundles (only for product-signal failures)

Skip this step entirely for a routine pass-rate report. Run it when the user asks why something failed, or when Step 2 found a test-phase failure worth root-causing:

bash
python3 .agents/skills/aicr-uat-report/uat_report.py --days 3 \
  --download-debug /tmp/uat-debug --max-downloads 3

Prefer --run <id> (repeatable) over raising --max-downloads: usually only the latest failure per failing combo is worth reading.

Bundles land in <dir>/<service>-<gpu>-<intent>-<run_id>/, each with a printed digest — MANIFEST head, failing checks from report.json, and a one-line contents summary. Read that summary for presence, not file names: a missing evidence/ or report.json says the run died before that stage, which is often the whole diagnosis. For a readiness-gate failure the digest also echoes the last --- failed validator output block from readiness-gate.log; its Failed resources: lines name the resources that never became ready, so start with the logs of the operator that owns them. Then open files per the Debug bundles table above.

Cite file_path:line and the run ID for every finding. Leave the download directory in place — the user may want to keep digging.

Step 3 — Render the report

Produce exactly two artifacts, in this order (see Output Format Reference). Sort the table worst-first: lowest pass ratio at the top; bold the Service/GPU/Intent cells of rows with product-signal failures. Include run URLs as links for at least the most recent failure of each failing combo.

Step 4 — Write the RC validation input

A numbered priority list derived from the table:

  1. Combos with consistent test-phase failures (0/N or repeated validate-phase or readiness-gate failures) — top manual-validation priority. For readiness failures, name the failing validators and the component that owns them.
  2. Combos whose most recent failure is test-phase (even if earlier ones were infra) — deserve a close look.
  3. Combos with infra/CI-only failures — noisy, not product signal; note them but rank low.
  4. Reservations in bring-up (e.g. GB200 while nightly-intents: [], kind lane) — expected churn, call out separately.
  5. Green combos — state them explicitly as lowest priority; a clean bill is information too.

Output Format Reference

markdown
## UAT report against `main` (<start>–<end>, N runs)

| Service | GPU | Intent | Pass | Failure nature |
|---|---|---|---|---|
| **AKS** | **H100** | **training** | **0/4** | Real test failures — every run fails at "UAT - validate (all phases)" ([latest](<url>)) |
| EKS | H100 | training | 2/5 | Infra only — 2x bring-up, 1x Buildx CI flake; no test-phase failures |
| GKE | H100 | training | 4/4 | Green |

## RC validation input

1. **<Service>/<GPU>/<intent> is the clear red flag** — <pass ratio,
   failure signature, latest run link, whether it also fails on release
   tags (env issue) or only main (regression candidate)>.
2. ...
N. **<green combos> are solidly green** — lowest manual-testing priority.

Keep failure-nature cells to one sentence; detail beyond that belongs in the RC list, not the table.

Failure Modes

  • gh run list returns nothing — window may predate retention or the workflow was renamed; say so rather than reporting "all green".
  • Unparsed titles in the Notes line — the run-name format in uat-run.yaml changed; read the workflow's current format string and update NEW_TITLE/OLD_TITLE in uat_report.py in the same PR.
  • Unknown reservation (e.g. new cloud) — the script falls back to the uppercased cloud token as the service name; cross-check new rows against infra/uat/reservations.yaml.
  • A combo has very few runs (e.g. 0/1) — flag low sample size instead of declaring it broken.
  • "no debug artifact" — expected for bring-up/Buildx/ingest failures (see Debug bundles). Report it as corroborating the infra classification; do not present it as a tooling problem.
  • "debug artifact expired" — the window exceeds the 30-day retention. Nothing to recover; note it and work from failed step names.
  • cluster-debug/ missing or thin — the collector is best-effort and its cloud credentials can expire on a long failure. Say the bundle is incomplete; do not read it as the cluster being healthy.
  • Readiness gate failed for an infra reason — the gate talks to the API server on every attempt, so expired cloud credentials or an unreachable control plane also fail it. If the failed-validator block is absent or its messages are connection/auth errors rather than resources that are not ready, reclassify to infra.
  • Failing step name disagrees with report.json — trust the bundle. UAT - validate (all phases) + emit signed evidence covers two concerns, so N/N passing checks under a failed step means the evidence leg failed, not the product. Reclassify to infra before ranking it in Step 4.

What This Skill Does NOT Do

  • Does not dispatch, re-run, or cancel workflow runs
  • Does not download raw job logs (gh run view --log); it reports failed step names and, on request, the uploaded cluster debug bundles
  • Does not modify reservations, workflows, or any in-repo file — the only writes are downloaded artifacts under the --download-debug directory

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/aicr-uat-report of NVIDIA/aicr.

  • SKILL.md
  • uat_report.py

Open the folder on GitHubat commit e8f18da

Compare with similar skills

Aicr Uat Report next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Aicr Uat Report compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Aicr Uat Report this skillNVIDIA/aicr440—~3.2kAutomated safety check: PassApache-2.0
Kcli Cluster Deploymentkarmab/kcli653—~1.5kAutomated safety check: PassApache-2.0
Apex Azure Cloud Migratejonathan-vella/apex217—~1.2kAutomated safety check: PassMIT
Devopsnicepkg/auto-company1952 repos~814Automated safety check: PassMIT
Provider Bug Reviewmondoohq/mql412—~2.9kAutomated safety check: PassCustom licence
Logfire Infrastructurepydantic/skills140—~1.8kAutomated safety check: PassMIT

Similar skills

  • Guides deployment and management of Kubernetes clusters with kcli.

    653 GitHub stars~1.5k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Apex Azure Cloud Migrate

    jonathan-vella/apex

    WORKFLOW SKILL — Assess and migrate cross-cloud workloads to Azure: assessments and code conversion from AWS, GCP, Heroku, Kubernetes or Spring.

    217 GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    195 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • Deep static code review of an mql provider for logic errors, nil-handling bugs, pagination truncation, caching/id collisions, and other defects that silently give users wrong data.

    412 GitHub stars~2.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Logfire Infrastructure

    pydantic/skills

    Official

    Monitor hosts, Docker containers, Kubernetes clusters, database/queue/cache servers, and cloud-provider metrics with Pydantic Logfire — no application code required.

    140 GitHub stars~1.8k tokensUpdated 9 days ago
    DevOps & CloudAuto-check passed
  • Extend Discovery Type

    runwhen-contrib/runwhen-local

    Add or enrich a resource type in an existing RunWhen Local discovery indexer (Azure azureapi, GCP gcpapi, AWS, or Kubernetes).

    163 GitHub stars~1.7k tokensUpdated today
    DevOps & CloudAuto-check passed

More from NVIDIA/aicr

All 10 skills in this repo
  • Official

    Multi-agent PR review using Claude Code, Codex, and CodeRabbit.

    440 GitHub stars~15k tokensUpdated today
    Auto-check passed
  • Official

    A skill your agent uses when analyzing an AICR snapshot YAML file, reviewing cluster state, comparing provider characteristics, extracting GPU/network topology insights, or generating a cluster…

    440 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Official

    Scaffolds an interactive guided demo script (demos/.sh), live or self-paced, with the Frame → Tell → Show → Close pattern.

    440 GitHub stars~929 tokensUpdated today
    Auto-check passed
  • Official

    A skill your agent uses when building a self-contained HTML slide deck or visual talking-point for a technical concept or workflow (e.g.

    440 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Official

    A skill your agent uses when drafting the human-readable GitHub release notes summary for an upcoming AICR release.

    440 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Official

    A skill your agent uses when reviewing the weekly AICR component drift report — the Slack digest and drift-report.json artifact produced by Registry Drift Report (registry-drift.yaml) listing which…

    440 GitHub stars~2.8k tokensUpdated today
    Auto-check passed

Categories

Questions about Aicr Uat Report

What does Aicr Uat Report do?

A skill your agent uses when reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB200) x intent combinations are passing or failing in the UAT Run…. Aicr Uat Report is an agent skill from NVIDIA/aicr, published by the product's own GitHub organization.yaml).

When should I use Aicr Uat Report?

Aicr Uat Report fits situations like: reporting on UAT health across services and GPU targets — which service (EKS/GKE/AKS) x GPU (H100/GB20; X intent combinations are passing; failing in the UAT Run workflow (uat-run.yaml); /aicr-uat-report.

How do I install Aicr Uat Report in Claude Code?

Run `npx skills add NVIDIA/aicr --skill aicr-uat-report -a claude-code`. Or copy the skill folder (.agents/skills/aicr-uat-report in NVIDIA/aicr) into .claude/skills/aicr-uat-report in your project. Claude Code loads it when a task matches its description.

How do I install Aicr Uat Report in Codex?

Run `npx skills add NVIDIA/aicr --skill aicr-uat-report -a codex`. Or copy the skill folder (.agents/skills/aicr-uat-report in NVIDIA/aicr) into .agents/skills/aicr-uat-report in your project. Codex loads it when a task matches its description.

Can I use Aicr Uat Report in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/aicr --skill aicr-uat-report -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/aicr-uat-report, .gemini/skills/aicr-uat-report, .github/skills/aicr-uat-report and .opencode/skills/aicr-uat-report in your project.

What does Aicr Uat Report need to run?

Going by SKILL.md and its folder, Aicr Uat Report needs Python for the scripts in its folder and the command-line tools its instructions call (gh and python3). Our summary lists: Python 3.

Does Aicr Uat Report access the network?

SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Aicr Uat Report safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Aicr Uat Report use?

Aicr Uat Report is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Aicr Uat Report use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Aicr Uat Report?

Skills that share tags, products or a category with Aicr Uat Report: Kcli Cluster Deployment (karmab/kcli, 653 stars), Apex Azure Cloud Migrate (jonathan-vella/apex, 217 stars), Devops (nicepkg/auto-company, 195 stars) and Provider Bug Review (mondoohq/mql, 412 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Aicr Uat Report?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/aicr, which has 440 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 10, 2026.

Source: NVIDIA/aicr on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.