Segment Anything Model Guide
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Sparse4D for multi-camera temporal 3D object detection and tracking.
$ npx skills add NVIDIA/skills --skill tao-train-sparse4d -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-train-sparse4d --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-sparse4d .claude/skills/tao-train-sparse4d && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-train-sparse4d" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-sparse4d into .claude/skills/tao-train-sparse4d/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-sparse4d", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-train-sparse4dType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-train-sparse4d -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-train-sparse4d --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-train-sparse4d .agents/skills/tao-train-sparse4d && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-train-sparse4d" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-sparse4d into .agents/skills/tao-train-sparse4d/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-sparse4d", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-sparse4d -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-train-sparse4d --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-train-sparse4d .cursor/skills/tao-train-sparse4d && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-train-sparse4d" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-sparse4d into .cursor/skills/tao-train-sparse4d/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-sparse4d", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-train-sparse4d--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-train-sparse4d -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-train-sparse4d --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-train-sparse4d .gemini/skills/tao-train-sparse4d && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-train-sparse4d" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-sparse4d into .gemini/skills/tao-train-sparse4d/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-sparse4d", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-train-sparse4dInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-train-sparse4d -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-train-sparse4d .github/skills/tao-train-sparse4d && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-sparse4d" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-sparse4d into .github/skills/tao-train-sparse4d/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-sparse4d", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-sparse4d -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-train-sparse4d --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-train-sparse4d .opencode/skills/tao-train-sparse4d && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-sparse4d" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-sparse4d into .opencode/skills/tao-train-sparse4d/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-sparse4d", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-train-sparse4dSparse4D for multi-camera temporal 3D object detection and tracking.
Tao Train Sparse4d is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Sparse4D for multi-camera temporal 3D object detection and tracking. Uses sparse queries with deformable attention across camera views and time for end-to-end 3D perception, with an instance bank for temporal tracking. Use when training, evaluating, exporting, quantizing, or running inference for a TAO Sparse4D model. Trigger phrases include "train Sparse4D", "multi-camera 3D detection", "temporal 3D tracker", "sparse query 3D perception".
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 27 other files, including scripts and reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.
It sits in AI & LLM Engineering, covering Computer vision. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires docker + nvidia-container-toolkit.
From compatibility in the SKILL.md frontmatter.
Tao Train Sparse4d loads about 3.8k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 116 tokens; SKILL.md has 1,238 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,238 words, ~3,762 tokens.
.claude/skills/tao-train-sparse4d/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Sparse4D for multi-camera temporal 3D object detection and tracking. Uses sparse queries with deformable attention across camera views and time for end-to-end 3D perception. Includes instance bank for temporal tracking.
Use a pretrained ResNet-101 backbone when one is available by setting
train.pretrained_model_path. For local smoke validation, Sparse4D training
can run with an empty train.pretrained_model_path, but production runs should
still use a compatible PTM.
Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML for an action requires schemas/<action>.schema.json and references/spec_template_<action>.yaml to exist and parse. Use the packaged selected-action schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.
This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.
Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.
img_bbox_NuScenes/mAP and mAP; AutoML metric
extractors should treat those emitted keys as aliases for val_mAP.
Multi-fidelity AutoML algorithms such as Hyperband, ASHA, and BOHB may
promote a checkpoint to a resume job that completes without emitting a fresh
val_mAP alias. In that case, compare AutoML's carried metric to the source
rung job that emitted img_bbox_NuScenes/mAP or mAP, while still verifying
that the promoted job resumed from the explicit epoch/step checkpoint,
produced a real checkpoint, and is usable for evaluate/inference.| Action | Spec Key | Source | Files | List? |
|---|---|---|---|---|
| dataset_convert | aicity.root | id | No | |
| evaluate | dataset.data_root | eval_dataset | (from convert job, spec: aicity.split) | No |
| evaluate | model.head.instance_bank.anchor | train_datasets | /results/{dataset_convert_job_id}/anchor_init.npy | No |
| evaluate | dataset.train_dataset.ann_file | train_datasets | (from convert job, spec: aicity.split) | No |
| evaluate | dataset.val_dataset.ann_file | eval_dataset | (from convert job, spec: aicity.split) | No |
| evaluate | dataset.test_dataset.ann_file | inference_dataset | (from convert job, spec: aicity.split) | No |
| export | model.head.instance_bank.anchor | train_datasets | /results/{dataset_convert_job_id}/anchor_init.npy | No |
| inference | dataset.data_root | inference_dataset | (from convert job, spec: aicity.split) | No |
| inference | model.head.instance_bank.anchor | train_datasets | /results/{dataset_convert_job_id}/anchor_init.npy | No |
| inference | dataset.train_dataset.ann_file | train_datasets | (from convert job, spec: aicity.split) | No |
| inference | dataset.val_dataset.ann_file | eval_dataset | (from convert job, spec: aicity.split) | No |
| inference | dataset.test_dataset.ann_file | inference_dataset | (from convert job, spec: aicity.split) | No |
| quantize | dataset.data_root | train_datasets | (from convert job, spec: aicity.split) | No |
| quantize | model.head.instance_bank.anchor | train_datasets | /results/{dataset_convert_job_id}/anchor_init.npy | No |
| quantize | dataset.train_dataset.ann_file | train_datasets | (from convert job, spec: aicity.split) | No |
| quantize | dataset.val_dataset.ann_file | eval_dataset | (from convert job, spec: aicity.split) | No |
| quantize | dataset.test_dataset.ann_file | inference_dataset | (from convert job, spec: aicity.split) | No |
| quantize | dataset.quant_calibration_dataset.images_dir | train_datasets | No | |
| train | dataset.data_root | train_datasets | (from convert job, spec: aicity.split) | No |
| train | model.head.instance_bank.anchor | train_datasets | /results/{dataset_convert_job_id}/anchor_init.npy | No |
| train | dataset.train_dataset.ann_file | train_datasets | (from convert job, spec: aicity.split) | No |
| train | dataset.val_dataset.ann_file | eval_dataset | (from convert job, spec: aicity.split) | No |
| train | dataset.test_dataset.ann_file | inference_dataset | (from convert job, spec: aicity.split) | No |
Data source overrides are mandatory for every action — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in spec_overrides.
S3_TRAIN = "s3://bucket/data/train"
S3_EVAL = "s3://bucket/data/eval"
CONVERTED_SCENE = "<scene-from-converter>" # e.g. "subsetscene+bev-sensor-random-0"train (mandatory data sources):
CONVERTED = "s3://bucket/results/<dataset_convert_job_id>"
{
"train.num_epochs": 30,
"train.checkpoint_interval": 10,
"train.validation_interval": 10,
"train.num_gpus": 1,
"dataset.sequences.split_num": 90,
"dataset.train_dataset.sequences_split_num": 90,
"dataset.data_root": f"{S3_TRAIN}/train",
"model.head.instance_bank.anchor": f"{CONVERTED}/anchor_init.npy",
"dataset.train_dataset.ann_file": f"{CONVERTED}/train/{CONVERTED_SCENE}_infos_train.pkl",
"dataset.val_dataset.ann_file": f"{CONVERTED}/val/{CONVERTED_SCENE}_infos_val.pkl",
"dataset.test_dataset.ann_file": f"{CONVERTED}/test/{CONVERTED_SCENE}_infos_test.pkl",
}evaluate (mandatory data sources):
CONVERTED = "s3://bucket/results/<dataset_convert_job_id>"
{
"dataset.data_root": f"{S3_EVAL}/val",
"model.head.instance_bank.anchor": f"{CONVERTED}/anchor_init.npy",
"dataset.train_dataset.ann_file": f"{CONVERTED}/train/{CONVERTED_SCENE}_infos_train.pkl",
"dataset.val_dataset.ann_file": f"{CONVERTED}/val/{CONVERTED_SCENE}_infos_val.pkl",
"dataset.test_dataset.ann_file": f"{CONVERTED}/test/{CONVERTED_SCENE}_infos_test.pkl",
}export (mandatory data sources):
CONVERTED = "s3://bucket/results/<dataset_convert_job_id>"
{
"model.head.instance_bank.anchor": f"{CONVERTED}/anchor_init.npy",
}inference (mandatory data sources):
CONVERTED = "s3://bucket/results/<dataset_convert_job_id>"
{
"dataset.data_root": f"{S3_EVAL}/test",
"model.head.instance_bank.anchor": f"{CONVERTED}/anchor_init.npy",
"dataset.train_dataset.ann_file": f"{CONVERTED}/train/{CONVERTED_SCENE}_infos_train.pkl",
"dataset.val_dataset.ann_file": f"{CONVERTED}/val/{CONVERTED_SCENE}_infos_val.pkl",
"dataset.test_dataset.ann_file": f"{CONVERTED}/test/{CONVERTED_SCENE}_infos_test.pkl",
}quantize (mandatory data sources):
CONVERTED = "s3://bucket/results/<dataset_convert_job_id>"
{
"dataset.data_root": f"{S3_TRAIN}/train",
"model.head.instance_bank.anchor": f"{CONVERTED}/anchor_init.npy",
"dataset.train_dataset.ann_file": f"{CONVERTED}/train/{CONVERTED_SCENE}_infos_train.pkl",
"dataset.val_dataset.ann_file": f"{CONVERTED}/val/{CONVERTED_SCENE}_infos_val.pkl",
"dataset.test_dataset.ann_file": f"{CONVERTED}/test/{CONVERTED_SCENE}_infos_test.pkl",
"dataset.quant_calibration_dataset.images_dir": f"{S3_TRAIN}",
}See references/local_docker_conversion.md for local-docker conversion roots and mounts, H5 depth-path normalization, converted annotation filenames, smoke-run max_num_cams/anchor contracts for export compatibility, and converted-artifact verification before train/evaluate/inference.
Optional. Val/test splits configured via dataset ann_file paths.
Launch method: Lightning-managed (single python process, Lightning spawns workers).
| Spec Key | Description | Default |
|---|---|---|
train.num_gpus | Number of GPUs | 1 |
train.gpu_ids | GPU device indices | [0] |
train.num_nodes | Number of nodes | 1 |
ddp_find_unused_parameters_true (no fsdp support)sync_batchnorm is always enabled (True)num_frames * num_bev_groups / (num_nodes * num_gpus * batch_size)Multi-node env vars (set by orchestrator): WORLD_SIZE, NODE_RANK, MASTER_ADDR, MASTER_PORT, NUM_GPU_PER_NODE.
Minimum 2 GPU(s), recommended 8 GPU(s). 40GB+ (A100 recommended) VRAM per GPU. Multi-camera temporal model is memory intensive. bf16 required for practical training. Multi-GPU strongly recommended. Instance bank requires substantial memory for temporal reasoning.
dataset_convert required: Must run dataset_convert first to produce annotation pickles and anchor_init.npy.
dataset_convert container/command: Sparse4D conversion is an AICity to
OVPKL annotations conversion. Launch dataset_convert with the action-level
tao_toolkit.data_services image and annotations convert -e {config_path};
do not use the PyTorch sparse4d CLI for conversion. Train/evaluate/export/
inference still use the model-level PyTorch image.
Stable raw-data path: The AICity to OVPKL converter writes image paths into
the generated pickle files. Keep aicity.root at /data/aicity_root during
conversion, then point dataset.data_root at the split folder, for example
/data/aicity_root/train for training or /data/aicity_root/val for
evaluation. This preserves the converter's absolute RGB paths and relative
depth paths.
H5 depth tuple mismatch: If training fails with an H5 path error where the
trainer tries to open a camera directory such as
/data/aicity_root/train/<scene>/Camera, run
models/tao-train-sparse4d/scripts/normalize_depth_paths.py --data-root <host-aicity-root>/train <converted-ann-dir>
after dataset_convert and before train/evaluate/inference. The helper rewrites
converted depth_map_path tuples to point at
<scene>/depth_maps/<camera>.h5 with the H5 dataset key basename.
Missing anchor file: Set model.head.instance_bank.anchor to the anchor_init.npy path from dataset_convert results.
Temporal OOM: Reduce dataset.num_frames or dataset.batch_size if running out of memory during temporal training.
Quantize image compatibility: The model-skill wiring should pass
quantize.model_path through the parent-model resolver, and checkpoint handoff
should select the exact epoch/step checkpoint just like evaluate, inference,
export, and resume. TorchAO checkpoint quantization passes in the
validation-fixes-20260525 PyT image and writes
quantized_model_torchao.pth. Older 7.0.0-rc PyT images may fail inside the
Sparse4D quantize entrypoint or lack ONNX quantization dependencies; do not
remove or skip the advertised quantize action if that occurs. Report the
container/image failure and keep the exact checkpoint path visible.
See references/spec_param_inference.md for the model-specific inference mappings from TAO Core sparse4d.config.json (the per-action spec-field to inference-function table) and the parent_model/parent_job_id checkpoint-resolution rules that generated runners apply with SDK helpers before create_job().
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 23 other files (scripts, references) in skills/tao-train-sparse4d of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Tao Train Sparse4d next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Train Sparse4d this skillNVIDIA/skills | 3.5k | — | ~3.8k | Automated safety check: Notes | Apache-2.0 | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | |
| CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Yolo Master AgentTencent/YOLO-Master | 742 | — | ~755 | Automated safety check: Pass | AGPL-3.0 | |
| Video Understandjjyaoao/HelloAgents | 3.2k | 1 repos | ~6.2k | Automated safety check: Pass | MIT | |
| Motioneyes Visual Analysisedwardsanchez/MotionEyes | 229 | — | ~2k | Automated safety check: Pass | None |
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
Orchestra-Research/AI-Research-SKILLs
Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.
Tencent/YOLO-Master
A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher.
jjyaoao/HelloAgents
Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.
edwardsanchez/MotionEyes
Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison.
Orchestra-Research/AI-Research-SKILLs
Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Categories
Sparse4D for multi-camera temporal 3D object detection and tracking. Tao Train Sparse4d is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Sparse4D for multi-camera temporal 3D object detection and tracking.
Tao Train Sparse4d fits situations like: running inference for a TAO Sparse4D model; phrases include train Sparse4D; multi-camera 3D detection; temporal 3D tracker.
Run `npx skills add NVIDIA/skills --skill tao-train-sparse4d -a claude-code`. Or copy the skill folder (skills/tao-train-sparse4d in NVIDIA/skills) into .claude/skills/tao-train-sparse4d in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-train-sparse4d -a codex`. Or copy the skill folder (skills/tao-train-sparse4d in NVIDIA/skills) into .agents/skills/tao-train-sparse4d in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-sparse4d -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-sparse4d, .gemini/skills/tao-train-sparse4d, .github/skills/tao-train-sparse4d and .opencode/skills/tao-train-sparse4d in your project.
SKILL.md names no scripts, command-line tools or credentials: Tao Train Sparse4d is instructions for the agent only. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Tao Train Sparse4d is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Train Sparse4d: Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Yolo Master Agent (Tencent/YOLO-Master, 742 stars) and Video Understand (jjyaoao/HelloAgents, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.