Megatron-LM CI Failure Triage
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
Performs deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation.
$ npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-analyze-changenet-rca --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-analyze-changenet-rca .claude/skills/tao-analyze-changenet-rca && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-analyze-changenet-rca" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-changenet-rca into .claude/skills/tao-analyze-changenet-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-changenet-rca", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-changenet-rcaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-analyze-changenet-rca --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-analyze-changenet-rca .agents/skills/tao-analyze-changenet-rca && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-analyze-changenet-rca" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-changenet-rca into .agents/skills/tao-analyze-changenet-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-changenet-rca", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-analyze-changenet-rca --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-analyze-changenet-rca .cursor/skills/tao-analyze-changenet-rca && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-analyze-changenet-rca" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-changenet-rca into .cursor/skills/tao-analyze-changenet-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-changenet-rca", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-analyze-changenet-rca--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-analyze-changenet-rca --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-analyze-changenet-rca .gemini/skills/tao-analyze-changenet-rca && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-analyze-changenet-rca" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-changenet-rca into .gemini/skills/tao-analyze-changenet-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-changenet-rca", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-analyze-changenet-rcaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-analyze-changenet-rca .github/skills/tao-analyze-changenet-rca && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-analyze-changenet-rca" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-changenet-rca into .github/skills/tao-analyze-changenet-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-changenet-rca", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-analyze-changenet-rca --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-analyze-changenet-rca .opencode/skills/tao-analyze-changenet-rca && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-analyze-changenet-rca" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-changenet-rca into .opencode/skills/tao-analyze-changenet-rca/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-changenet-rca", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-analyze-changenet-rcaPerforms deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation.
Tao Analyze Changenet Rca is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Performs deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation. Use when analyzing ChangeNet model failures, investigating poor recall / FAR / PASS-NOPASS metrics, auditing visual inspection pipeline quality, or running an RCA report for an AOI defect-detection model. Trigger phrases include "RCA on my ChangeNet model", "why is my AOI model failing", "audit ChangeNet predictions", "investigate FAR regressions", "root cause analysis on…
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit. Workflows declare additional requirements.
It sits in Development, covering Root cause analysis. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires docker + nvidia-container-toolkit. Workflows declare additional requirements.
From compatibility in the SKILL.md frontmatter.
Tao Analyze Changenet Rca loads about 1.6k tokens when it runs, and up to ~9.6k if it reads all its reference files. Until then it costs about 140 tokens; SKILL.md has 731 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 731 words, ~1,616 tokens.
.claude/skills/tao-analyze-changenet-rca/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
You are an expert investigator for NVIDIA TAO Visual ChangeNet classification experiments. Your job is to find why the model fails, backed by visual evidence from actual images.
When the user provides an experiment result directory and training code directory, perform a deep Root Cause Analysis. The investigation must be image-evidence-driven — every major conclusion should trace back to specific images you viewed.
train/ and inference/visual_changenet/ source treeThe ChangeNet model compares a test image against a golden image (known-good reference) to detect differences. When viewing images, check these three things:
The investigation has 5 phases. Phase 1 (numbers) gives you hypotheses. Phase 2 (images) proves or disproves them. Phase 3 (cross-dimensional) finds hidden patterns. Phase 4 (config) explains the mechanism. Phase 5 (counterfactual) quantifies fixes. Phase 2 is the core — spend the most effort there. Phase 5 is the most actionable — never skip it.
See references/investigation-phases.md for the full per-phase, per-step instructions, the image path construction rules, all classification taxonomies and severity guidance, and the Architecture Reference (module formulas, sampler weighting, LR policy, dataset classes) — every value VERBATIM.
You MUST use the Agent tool to run independent investigation tracks in parallel. Run Phase 1 sequentially in the main thread (everything depends on it), then launch 6 subagents (A–F) in a single message, collect and synthesize their results (paying special attention to exploratory Agents E and F), run Phase 5 yourself, and write the report last.
Before writing RCA_Report.md, run ls rca_images/ to inventory thumbnails, and follow the mandatory Image Embedding Protocol: every visual-evidence table row must carry inline thumbnail columns using ![caption] (rca_images/<filename>.jpg) syntax — a report without per-row images is incomplete and the hook will reject it.
See references/parallelization.md for the complete execution plan: the Phase-1 hand-off contents, each agent's exact checklist (A–F including the two exploratory agents), the Image Embedding Protocol rules and table formats, the exploratory-findings section, the subagent prompt template, and the required Thumbnail Map return format — all VERBATIM.
Produce RCA_Report.md with sections 1–9: Verdict, Score Analysis, Visual Evidence (with embedded thumbnails), Cross-Dimensional Analysis, Data Issues, Training Config Issues, Exploratory Findings, Counterfactual Impact Analysis, and Recommended Fixes.
Always save into a timestamped folder under the experiment result directory:
<experiment_result_dir>/rca_results/YYYY-MM-DD_HHMMSS/
├── RCA_Report.md
├── rca_images/
├── rca_config/
└── claude_session.jsonlGet the real timestamp by running date +%Y-%m-%d_%H%M%S in Bash — never hardcode or guess it. If the user specifies a custom path, use that instead but keep the same structure.
See references/output-structure.md for the complete section-by-section report skeleton (every table header and summary line) and the full output layout with hook-copied contents — VERBATIM.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 15 other files (references) in skills/tao-analyze-changenet-rca of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Tao Analyze Changenet Rca next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Analyze Changenet Rca this skillNVIDIA/skills | 3.6k | — | ~1.6k | Automated safety check: Notes | Apache-2.0 | |
| Megatron-LM CI Failure TriageNVIDIA/Megatron-LM | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Analysissgl-project/sglang | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| OpenLogi macOS Permissions TriageAprilNEA/OpenLogi | 23k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Bug Finder for daisyUIsaadeghi/daisyui | 43k | — | ~2.3k | Automated safety check: Pass | MIT | |
| Root Cause Debugginggarrytan/gstack | 136k | — | ~1.4k | Automated safety check: Pass | MIT |
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
AprilNEA/OpenLogi
Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.
saadeghi/daisyui
Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.
garrytan/gstack
Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.
apache/shardingsphere
Review Apache ShardingSphere or user-authorized downstream pull requests and PR discussions from public or authorized repository evidence.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Performs deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation. Tao Analyze Changenet Rca is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Performs deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation.
Tao Analyze Changenet Rca fits situations like: analyzing ChangeNet model failures; investigating poor recall / FAR / PASS-NOPASS metrics; auditing visual inspection pipeline quality; running an RCA report for an AOI defect-detection model.
Run `npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a claude-code`. Or copy the skill folder (skills/tao-analyze-changenet-rca in NVIDIA/skills) into .claude/skills/tao-analyze-changenet-rca in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a codex`. Or copy the skill folder (skills/tao-analyze-changenet-rca in NVIDIA/skills) into .agents/skills/tao-analyze-changenet-rca in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-analyze-changenet-rca -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-analyze-changenet-rca, .gemini/skills/tao-analyze-changenet-rca, .github/skills/tao-analyze-changenet-rca and .opencode/skills/tao-analyze-changenet-rca in your project.
Going by SKILL.md and its folder, Tao Analyze Changenet Rca needs a shell for the scripts in its folder. Our summary lists: A Bash shell; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit. Workflows declare additional requirements..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Analyze Changenet Rca is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Analyze Changenet Rca: Megatron-LM CI Failure Triage (NVIDIA/Megatron-LM, 18k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars) and Bug Finder for daisyUI (saadeghi/daisyui, 43k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.