Official agent skill

Tao Train Visual Changenet

by NVIDIA in NVIDIA/skills

Visual ChangeNet for binary image classification and segmentation in AOI defect detection.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Train Visual Changenet

skills CLI
$ npx skills add NVIDIA/skills --skill tao-train-visual-changenet -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-train-visual-changenet --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-visual-changenet .claude/skills/tao-train-visual-changenet && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-train-visual-changenet
GitHub stars
3.5k
Token cost
~4.7k tokens
SKILL.md length
1,632 words
Files
44 (incl. scripts, references)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Visual ChangeNet for binary image classification and segmentation in AOI defect detection.

  • Running inference for PCB defect detection
  • SKILL.md covers Dataclass Schemas, Train Action Policy, Training Requirements and Running via local Docker, plus 7 more sections
  • Calls python3 and docker
  • Visual inspection

What it does

Tao Train Visual Changenet is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Visual ChangeNet for binary image classification and segmentation in AOI defect detection. Use when training, evaluating, exporting, or running inference for PCB defect detection or visual inspection, comparing image pairs for PASS/NOPASS classification, or producing change-segmentation masks. Trigger phrases include "train Visual ChangeNet", "ChangeNet classify", "ChangeNet segment", "AOI defect detection", "PCB inspection model".

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 46 other files, including scripts and reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.

It sits in AI & LLM Engineering, covering Computer vision. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Running inference for PCB defect detection
  • Visual inspection
  • Comparing image pairs for PASS/NOPASS classification
  • Producing change-segmentation masks

Example prompts

  • “train Visual ChangeNet”
  • “ChangeNet classify”
  • “ChangeNet segment”
  • “/tao-train-visual-changenet”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit.
  • Pre-approved tools (allowed-tools): Read, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Train Visual Changenet loads about 4.7k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 1,632 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,632 words, ~4,700 tokens.

Download SKILL.mdSave it as .claude/skills/tao-train-visual-changenet/SKILL.md (or your agent's skills folder). This skill also uses 43 other files; get the full folder from GitHub.
name
tao-train-visual-changenet
description
Visual ChangeNet for binary image classification and segmentation in AOI defect detection. Use when training, evaluating, exporting, or running inference for PCB defect detection or visual inspection, comparing image pairs for PASS/NO_PASS classification, or producing change-segmentation masks. Trigger phrases include "train Visual ChangeNet", "ChangeNet classify", "ChangeNet segment", "AOI defect detection", "PCB inspection model".
allowed-tools
Read, Bash
compatibility
Requires docker + nvidia-container-toolkit.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
pcb, aoi, defect, classification, segmentation, siamese, visual-inspection

Visual ChangeNet

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Visual ChangeNet is a TAO Toolkit model for visual inspection and defect detection. It supports two tasks:

  • Classify — Binary image classification using a siamese-style architecture with a shared backbone (C-RADIO ViT) and a learnable difference module. Compares image pairs to classify defects as PASS/NO_PASS.
  • Segment — Pixel-level change segmentation using a ViT-Large NVDINOv2 backbone. Compares before/after image pairs to produce a binary change mask.

Classify supports the public C-RADIOv2-B backbone and six frozen DINOv3 variants. Read references/dinov3-backbones.md before selecting DINOv3; it contains the exact variant map, freeze requirement, Hugging Face access rules, and local-staging overlay. For C-RADIO, use the bundled scripts/stage_backbone.py and the mount in references/local-docker.md.

Segment specs use model.backbone.type: vit_large_nvdinov2 and the NVDINOv2 checkpoint family. Keep the checkpoint architecture aligned with the backbone type: NV_DINOV2_518_16_256.ckpt is compatible with the packaged segment templates, but it must not be used with fan_small_12_p4_hybrid. If you switch to a different segment backbone, use a matching checkpoint or leave model.backbone.pretrained_backbone_path empty for default initialization.

Dataclass Schemas

Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML still requires schemas/train.schema.json and references/spec_template_train.yaml to exist and parse. Use the packaged train schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.

Train Action Policy

This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.

Checkpoint retention is an orchestration policy, not an HPO parameter. Both packaged train templates default train.checkpointer.enable_topk and train.checkpointer.replace_periodic to false, preserving periodic saves controlled by train.checkpoint_interval. When AutoML checkpoint retention is enabled, the AutoML runner sets both flags to true, monitors val_loss in min mode, and uses save_top_k: 1; this replaces the periodic series with the single best checkpoint. When AutoML checkpoint retention is disabled, leave the bounded-retention overrides unset so periodic checkpoint behavior remains.

Non-train actions declared by this model skill (evaluate, inference, export, quantize, segment_evaluate, and segment_inference) stay in this model skill. Do not present segment_export or segment_quantize as runnable parent-skill actions until matching entries are packaged in schemas/manifest.json. Prune and retrain are not declared in the current parent references/skill_info.yaml; do not present them as runnable parent-skill actions unless the metadata is extended with matching action wiring and schemas. The per-run automl_policy override does not change model metadata.

For TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and TensorRT inference for classify and segment variants), read references/tao-deploy-visual-changenet.md first. Deploy spec templates live in this skill's references/ folder with the spec_template_deploy_*.yaml prefix. Deploy requires an exported ONNX artifact as parent_model. If no ONNX artifact exists and the main skill does not expose an export action, report deploy as blocked instead of inventing an artifact.

Training Requirements

Visual ChangeNet has two separate task modes with different dataset types and data source structures.

Classify
  • Dataset type: visual_changenet_classify
  • Formats: default
  • Accepted dataset intents: training, evaluation, testing, calibration
  • Monitoring metric: val_loss
  • Standalone evaluate metric: test_acc (with test_fpr, test_fnr, and defect_acc also emitted). AutoML must rank recommendations by the training val_loss; use test_acc only to verify that the selected checkpoint loads and evaluates successfully. Do not expect the evaluate action to emit val_loss.
Per-Action Dataset Requirements (Classify)

The quantize and gen_trt_engine rows below describe TAO spec data requirements only. They are not parent-skill actions unless the corresponding action is declared in references/skill_info.yaml or deploy/skill_info.yaml.

ActionSpec KeySourceFilesList?
traindataset.classify.train_dataset.images_dirtrain_datasetsimages.tar.gzNo
traindataset.classify.train_dataset.csv_pathtrain_datasetsdataset.csvNo
traindataset.classify.validation_dataset.images_direval_datasetimages.tar.gzNo
traindataset.classify.validation_dataset.csv_patheval_datasetdataset.csvNo
quantizedataset.classify.train_dataset.images_dirtrain_datasetsimages.tar.gzNo
quantizedataset.classify.train_dataset.csv_pathtrain_datasetsdataset.csvNo
quantizedataset.classify.validation_dataset.images_direval_datasetimages.tar.gzNo
quantizedataset.classify.validation_dataset.csv_patheval_datasetdataset.csvNo
quantizedataset.classify.quant_calibration_dataset.images_dirtrain_datasetsimages.tar.gzNo
evaluatedataset.classify.validation_dataset.images_direval_datasetimages.tar.gzNo
evaluatedataset.classify.validation_dataset.csv_patheval_datasetdataset.csvNo
evaluatedataset.classify.test_dataset.images_direval_datasetimages.tar.gzNo
evaluatedataset.classify.test_dataset.csv_patheval_datasetdataset.csvNo
inferencedataset.classify.infer_dataset.images_dirinference_datasetimages.tar.gzNo
inferencedataset.classify.infer_dataset.csv_pathinference_datasetdataset.csvNo
gen_trt_enginegen_trt_engine.tensorrt.calibration.cal_image_dircalibration_datasetimages.tar.gzYes
Segment
  • Dataset type: visual_changenet_segment
  • Formats: default
  • Accepted dataset intents: training, calibration
  • Monitoring metric: val_loss

Segment uses a paired directory structure (A/, B/, list/, label/) instead of CSV + images. The root_dir spec key points to the top-level directory containing all four subdirectories.

Required files per dataset: A.tar.gz, B.tar.gz, list.tar.gz, label.tar.gz

Per-Action Dataset Requirements (Segment)

The quantize and gen_trt_engine rows below describe TAO spec data requirements only. They are not parent-skill actions unless the corresponding action is declared in references/skill_info.yaml or deploy/skill_info.yaml.

ActionSpec KeySourceFilesList?
traindataset.segment.root_dirtrain_datasets(root directory)No
quantizedataset.segment.root_dirtrain_datasets(root directory)No
quantizedataset.segment.quant_calibration_dataset.images_dirtrain_datasets(root directory)No
evaluatedataset.segment.root_dirtrain_datasets(root directory)No
inferencedataset.segment.root_dirtrain_datasets(root directory)No
gen_trt_enginedataset.segment.root_dirtrain_datasets(root directory)No
gen_trt_enginegen_trt_engine.tensorrt.calibration.cal_image_dircalibration_datasetimages.tar.gzYes
Show full SKILL.md (691 more words)Show less
Typical Spec Overrides

Data source overrides are mandatory for every action — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in spec_overrides.

python
S3_TRAIN = "s3://bucket/data/train"
S3_EVAL = "s3://bucket/data/eval"

train (classify, mandatory data sources):

python
{
    "train.num_epochs": 30,
    "train.checkpoint_interval": 10,
    "train.validation_interval": 10,
    "train.num_gpus": 1,
    "train.use_distributed_sampler": False,
    "train.sync_batchnorm": False,
    "dataset.classify.train_dataset.images_dir": f"{S3_TRAIN}/images.tar.gz",
    "dataset.classify.train_dataset.csv_path": f"{S3_TRAIN}/dataset.csv",
    "dataset.classify.validation_dataset.images_dir": f"{S3_EVAL}/images.tar.gz",
    "dataset.classify.validation_dataset.csv_path": f"{S3_EVAL}/dataset.csv",
}

train (segment, mandatory data sources):

python
{
    "train.num_epochs": 30,
    "train.checkpoint_interval": 10,
    "train.validation_interval": 10,
    "train.num_gpus": 1,
    "train.use_distributed_sampler": False,
    "train.sync_batchnorm": False,
    "dataset.segment.root_dir": f"{S3_TRAIN}",
}

export (classify):

python
{
    "export.input_height": 896,
    "export.input_width": 224,
}

export (segment):

python
{
    "export.input_height": 224,
    "export.input_width": 224,
}

quantize (classify, mandatory data sources):

python
{
    "dataset.classify.train_dataset.images_dir": f"{S3_TRAIN}/images.tar.gz",
    "dataset.classify.train_dataset.csv_path": f"{S3_TRAIN}/dataset.csv",
    "dataset.classify.validation_dataset.images_dir": f"{S3_EVAL}/images.tar.gz",
    "dataset.classify.validation_dataset.csv_path": f"{S3_EVAL}/dataset.csv",
    "dataset.classify.quant_calibration_dataset.images_dir": f"{S3_TRAIN}/images.tar.gz",
}

evaluate (classify, mandatory data sources):

python
{
    "dataset.classify.validation_dataset.images_dir": f"{S3_EVAL}/images.tar.gz",
    "dataset.classify.validation_dataset.csv_path": f"{S3_EVAL}/dataset.csv",
    "dataset.classify.test_dataset.images_dir": f"{S3_EVAL}/images.tar.gz",
    "dataset.classify.test_dataset.csv_path": f"{S3_EVAL}/dataset.csv",
}

inference (classify, mandatory data sources):

python
{
    "dataset.classify.infer_dataset.images_dir": f"{S3_EVAL}/images.tar.gz",
    "dataset.classify.infer_dataset.csv_path": f"{S3_EVAL}/dataset.csv",
}

gen_trt_engine (classify, mandatory data sources):

python
{
    "gen_trt_engine.tensorrt.calibration.cal_image_dir": [f"{S3_TRAIN}/images.tar.gz"],
}

quantize (segment, mandatory data sources):

python
{
    "dataset.segment.root_dir": f"{S3_TRAIN}",
    "dataset.segment.quant_calibration_dataset.images_dir": f"{S3_TRAIN}",
}

evaluate (segment, mandatory data sources):

python
{
    "dataset.segment.root_dir": f"{S3_TRAIN}",
}

inference (segment, mandatory data sources):

python
{
    "dataset.segment.root_dir": f"{S3_TRAIN}",
}

gen_trt_engine (segment, mandatory data sources):

python
{
    "dataset.segment.root_dir": f"{S3_TRAIN}",
    "gen_trt_engine.tensorrt.calibration.cal_image_dir": [f"{S3_TRAIN}/images.tar.gz"],
}

Running via local Docker

Use the pinned TAO pyt image and invoke visual_changenet <train|evaluate|inference|export|quantize> directly. --shm-size=8g is required, the C-RADIO .safetensors must be mounted to /data/pretrained_models/C-RADIOv2_B.safetensors, and checkpoint/results_dir can be overridden on the command line. See references/local-docker.md for the full docker run command, mounts, and overrides.

Tasks

Classify (default)

Uses actions: train, evaluate, inference. Defaults template: references/spec_template_train.yaml. evaluate / inference need a checkpoint from a prior 7.1 train under results_dir — there is no pretrained 7.1 classify checkpoint on NGC (the 7.0-era visual_changenet_nvpcb_trainable_v1.0 fails to load on 7.1 with a radio.* KeyError). On a fresh workspace, train from the public backbone first; do not try to download an NGC full_model classify checkpoint, and do not hardcode an NGC org.

Segment

Uses skill action names segment_train, segment_evaluate, and segment_inference. When invoking local Docker directly, run TAO CLI subcommands train, evaluate, and inference with task: segment in the spec. The schema-driven action templates are references/spec_template_segment_train.yaml, references/spec_template_segment_evaluate.yaml, and references/spec_template_segment_inference.yaml; the compact direct-Docker example template is references/spec_template_segment.yaml.

Segmentation requires compiling custom CUDA ops (MultiScaleDeformableAttention) on first run, which takes ~5 minutes. The ViT adapter backbone uses these for multi-scale feature extraction.

Dataset structure for segmentation differs from classify — uses paired directories (A/, B/, list/, label/) instead of CSV files. See dataset.segment.root_dir in the defaults.

Data Format

Classify needs a 4-column CSV (input_path,golden_path,label,object_name) plus an images directory; segment uses a paired directory structure (A/, B/, list/, label/) under dataset.segment.root_dir instead of CSV. The image_ext field (default .jpg) must match the actual file extensions; if images are .png, set dataset.classify.image_ext: .png. Multi-lighting input is configured via dataset.classify.input_map (each lighting name maps to a channel index) with dataset.classify.num_input set to match. See references/data-formats.md for the per-field input tables (classify train/eval/inference, segment), CSV column semantics, lighting/path-concatenation conventions, the segment directory layout, and input_map/grid_map examples.

Pre-Flight: validate the classify dataset (mandatory before every classify run)

Before launching a classify train, evaluate, or inference job, validate the CSV so a malformed dataset fails in <1s on the host instead of minutes into the GPU container (or, for a single-class train set, only after a checkpoint is written). Run:

bash
python3 skills/models/tao-train-visual-changenet/scripts/validate_vcn_dataset.py \
  --csv        <abs path to dataset.csv> \
  --images-dir <abs path to images dir> \
  --mode       train \
  --batch-size <dataset.classify.batch_size> --num-gpus <train.num_gpus>
# --mode: train | evaluate | inference

Exit 0 → launch. Exit 2 → fix the dataset, do not launch. The script rejects absolute CSV paths, flat filenames where a per-sample directory is required, single-class training sets, and a batch larger than the dataset. See references/data-formats.md for the per-check contract and the --light / --image-ext options.

Important Parameters

Key knobs include train.validation_interval (default 50, must be ≤ num_epochs), train.checkpoint_interval (default 200, must be ≤ num_epochs when periodic checkpointing is active), train.num_epochs (default 100), model.classify.eval_margin (default 0.3, the precision/recall threshold), model.classify.train_margin_euclid (default 2.0), model.classify.embedding_vectors (default 5), dataset.classify.batch_size (default 16, must be > 1), dataset.classify.fpratio_sampling (default 0.25), and train.classify.cls_weight (default [1.0, 10.0]). The train.checkpointer fields are fixed lifecycle controls, not HPO search parameters. Hardware: minimum 1 GPU with 16GB+ VRAM, recommended 8 GPUs (DDP); do not set gpu_spec_key (GPU count is managed internally by TAO), num_nodes (default 1) controls multi-node. See references/tuning-parameters.md for the full per-parameter guidance and hardware detail.

Error Patterns

For checkpoint-not-found, CSV format mismatch, image extension mismatch, OOM, low evaluation accuracy, the contrastive-loss AssertionError, checkpoint load key mismatch at evaluate/inference, non-convergence, segment-only backbone dimension mismatch, the MultiScaleDeformableAttention OSError, the Lightning MisconfigurationException, ModuleNotFoundError: nvidia_tao_pytorch, and epoch defaults, see references/troubleshooting.md for the full symptom-and-fix list.

Spec Param / Parent Model Inference

Model-specific parent-model mappings are declared in references/skill_info.yaml under spec_params, so agents resolve checkpoints before launching a job instead of guessing file names. For parent_model or parent_model_folder, pass the upstream train/export/AutoML child job id as the parent job id; list the parent result folder, filter checkpoint artifacts, and select the resolved model file or folder. See references/parent-model-inference.md for the full per-action spec-field-to-inference-function mapping table.

Deployment

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 43 other files (scripts, references) in skills/tao-train-visual-changenet of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • eval.config
  • evals/evals.json
  • references/data-formats.md
  • references/dinov3-backbones.md
  • references/local-docker.md
  • references/parent-model-inference.md
  • references/skill_info.yaml
  • references/spec_template_deploy_classify_evaluate.yaml
  • references/spec_template_deploy_classify_gen_trt_engine.yaml
  • references/spec_template_deploy_classify_inference.yaml
  • references/spec_template_deploy_segment_evaluate.yaml
  • references/spec_template_deploy_segment_gen_trt_engine.yaml
  • references/spec_template_deploy_segment_inference.yaml
  • references/spec_template_evaluate.yaml
  • references/spec_template_export.yaml
  • … and 26 more

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Tao Train Visual Changenet next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Train Visual Changenet compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Train Visual Changenet this skillNVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Yolo Master AgentTencent/YOLO-Master745—~755Automated safety check: PassAGPL-3.0
Video Understandjjyaoao/HelloAgents3.2k1 repos~6.2kAutomated safety check: PassMIT
LLaVA Vision-Language ModelOrchestra-Research/AI-Research-SKILLs13k6 repos~2kAutomated safety check: PassMIT

Similar skills

  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Yolo Master Agent

    Tencent/YOLO-Master

    A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher.

    745 GitHub stars~755 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Video Understand

    jjyaoao/HelloAgents

    Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.

    3.2k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • LLaVA Vision-Language Model

    Orchestra-Research/AI-Research-SKILLs

    Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.

    13k GitHub starsUsed in 6 repos~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Motioneyes Visual Analysis

    edwardsanchez/MotionEyes

    Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison.

    229 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Tao Train Visual Changenet

What does Tao Train Visual Changenet do?

Visual ChangeNet for binary image classification and segmentation in AOI defect detection. Tao Train Visual Changenet is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Visual ChangeNet for binary image classification and segmentation in AOI defect detection.

When should I use Tao Train Visual Changenet?

Tao Train Visual Changenet fits situations like: running inference for PCB defect detection; visual inspection; comparing image pairs for PASS/NOPASS classification; producing change-segmentation masks.

How do I install Tao Train Visual Changenet in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-train-visual-changenet -a claude-code`. Or copy the skill folder (skills/tao-train-visual-changenet in NVIDIA/skills) into .claude/skills/tao-train-visual-changenet in your project. Claude Code loads it when a task matches its description.

How do I install Tao Train Visual Changenet in Codex?

Run `npx skills add NVIDIA/skills --skill tao-train-visual-changenet -a codex`. Or copy the skill folder (skills/tao-train-visual-changenet in NVIDIA/skills) into .agents/skills/tao-train-visual-changenet in your project. Codex loads it when a task matches its description.

Can I use Tao Train Visual Changenet in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-visual-changenet -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-visual-changenet, .gemini/skills/tao-train-visual-changenet, .github/skills/tao-train-visual-changenet and .opencode/skills/tao-train-visual-changenet in your project.

What does Tao Train Visual Changenet need to run?

Going by SKILL.md and its folder, Tao Train Visual Changenet needs the command-line tools its instructions call (python3 and docker). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..

Does Tao Train Visual Changenet access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tao Train Visual Changenet safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Tao Train Visual Changenet use?

Tao Train Visual Changenet is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Train Visual Changenet use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.

What are the alternatives to Tao Train Visual Changenet?

Skills that share tags, products or a category with Tao Train Visual Changenet: Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Yolo Master Agent (Tencent/YOLO-Master, 745 stars) and Video Understand (jjyaoao/HelloAgents, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Train Visual Changenet?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.