Yolo Detection 2026
SharpAI/DeepCamera
YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.
DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection.
$ npx skills add NVIDIA/skills --skill tao-train-dino -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-train-dino --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-dino .claude/skills/tao-train-dino && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-train-dino" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-dino into .claude/skills/tao-train-dino/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-dino", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-train-dinoType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-train-dino -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-train-dino --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-train-dino .agents/skills/tao-train-dino && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-train-dino" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-dino into .agents/skills/tao-train-dino/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-dino", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-dino -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-train-dino --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-train-dino .cursor/skills/tao-train-dino && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-train-dino" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-dino into .cursor/skills/tao-train-dino/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-dino", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-train-dino--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-train-dino -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-train-dino --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-train-dino .gemini/skills/tao-train-dino && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-train-dino" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-dino into .gemini/skills/tao-train-dino/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-dino", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-train-dinoInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-train-dino -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-train-dino .github/skills/tao-train-dino && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-dino" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-dino into .github/skills/tao-train-dino/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-dino", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-dino -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-train-dino --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-train-dino .opencode/skills/tao-train-dino && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-dino" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-dino into .opencode/skills/tao-train-dino/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-dino", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-train-dinoDINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection.
Tao Train Dino is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Transformer-based detector with denoising training, multi-scale features, and optional distillation support. Use when training, evaluating, exporting, distilling, quantizing, or running inference for a TAO DINO detector. Trigger phrases include "train DINO", "DETR object detection", "TAO 2D detection", "DINO with distillation".
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 34 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.
It sits in AI & LLM Engineering, covering Computer vision. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires docker + nvidia-container-toolkit.
From compatibility in the SKILL.md frontmatter.
Tao Train Dino loads about 2.8k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 105 tokens; SKILL.md has 1,194 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,194 words, ~2,790 tokens.
.claude/skills/tao-train-dino/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Transformer-based detector with denoising training, multi-scale features, and optional distillation support.
Uses pretrained backbone weights (e.g. ResNet-50 ImageNet). Set model.pretrained_backbone_path for backbone-only or train.pretrained_model_path for full model.
Train, evaluate, export, distill, quantize, or run inference for a TAO DINO 2D object detector.
For TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and
TensorRT inference), read references/tao-deploy-dino.md first. Deploy spec templates live
in this skill's references/ folder with the spec_template_deploy_*.yaml
prefix.
references/dino-data-specs.md — dataset contracts, per-action dataset requirements, per-action spec-override examples (train, evaluate, export, deploy/gen_trt_engine, inference, quantize, distill), data-source arrays, checkpoint inference, and dataset layout.references/dino-actions-errors.md — important parameters, default values, evaluate/export defaults, hardware, and the full error-pattern catalog.references/dino-tuning-multigpu.md — full AutoML/HPO notes (metrics, hyperparameters, extractor) and multi-GPU spec consistency.references/tao-deploy-dino.md — TensorRT deploy workflow.references/detailed-guide.md — map to the detailed model guide.Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML still requires schemas/train.schema.json and references/spec_template_train.yaml to exist and parse. Use the packaged train schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.
This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.
Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.
The agent MUST read this section before generating any training or AutoML script for DINO.
test_mAP50 with maximize direction. Use test_mAP only when the user explicitly requests COCO/paper-style mAP.val_mAP50 (logged as Validation mAP50) for quick operational checks; val_mAP for COCO/paper-style benchmark comparisons.test_mAP50 for AP50 and test_mAP for COCO mAP. AutoML workflows that score recommendations with the standalone evaluate action must use the corresponding test_* KPI.Required datasets — MUST resolve both:
| Dataset | Required | Why |
|---|---|---|
| Train dataset URI | Yes | Training data (COCO format) |
| Validation dataset URI | Yes — ALWAYS | DINO unconditionally builds a val dataloader. Omitting val_data_sources causes FileNotFoundError at startup regardless of the metric or workflow. If the user has no separate eval split, reuse the train URI. |
Required inputs before generating any training spec:
num_classes — How many object classes? Default 91 (COCO). Must be >= max(category_id) + 1. Too low causes CUDA error: device-side assert triggered.Resolve these from the user request or the default profile below. Prompt only for values that are still missing after applying the profile rules.
Bankable local default profile for DINO AutoML smoke runs:
Use this profile only when the user asks to run DINO AutoML and does not provide dataset or class-count inputs. This profile is intentionally small and local to this skill bank; it is for smoke/iteration runs, not a production benchmark. Do not search previous runners, logs, session state, shell history, or the home directory to recover these values.
DINO_AUTOML_PROFILE = {
"train_dataset_uri": "s3://nvcf-storage-handling/data/tao_od_synthetic_subset_train_no_convert",
"validation_dataset_uri": "s3://nvcf-storage-handling/data/tao_od_synthetic_subset_val_no_convert",
"object_classes": 4,
"dataset_num_classes": 5,
"image_archive": "images.tar.gz",
"annotation_file": "annotations.json",
"max_recommendations": 10,
"train_num_epochs": 10,
"train_checkpoint_interval": 10,
"train_validation_interval": 1,
"train_num_gpus": 1,
}If the user supplies any dataset URI or class-count value, prefer the user value and ask for any remaining required DINO value. Do not partially mix a user's custom dataset with this profile's class count unless the user confirms it.
Do not prompt for image layout for the standard DINO dataset. The standard
TAO DINO dataset artifact is images.tar.gz plus annotations.json. Use
images.tar.gz in the remote image_dir spec override. The SDK downloads the
archive and rewrites the runtime spec to the extracted folder named after the
archive stem (images.tar.gz -> images). Only deviate if the user explicitly
provides a different image artifact name.
DINO supports train, evaluate, export, distill, quantize, and inference. Data-source
overrides are mandatory for every action — DINO's config.json has empty
data_sources because the runner cannot auto-resolve array-of-objects spec keys.
The agent MUST construct data source paths and include them in spec_overrides.
See references/dino-data-specs.md for the per-action dataset requirements table,
the standard dataset artifact (images.tar.gz + annotations.json) and runtime
folder rewrite rules, and the complete per-action spec_overrides examples for
train, evaluate, export, deploy/gen_trt_engine, inference, quantize, and distill —
including checkpoint inference via parent_model, the results_dir/train/
checkpoint location, and the distillation FAN-teacher / student rules.
Default values: num_epochs=10, batch_size=4, learning_rate=2e-4,
lr_backbone=2e-5, num_classes=91, backbone=resnet_50.
max(category_id) + 1. Too low causes CUDA error: device-side assert triggered. Set as <num_classes> + 1 in spec overrides.See references/dino-actions-errors.md for the full parameter list (backbone
options, train.optim.lr/lr_steps, model.num_queries, batch_size),
default values, evaluate defaults, export defaults (input 960x544, opset 17,
TRT data types, workspace 1024 MB), and hardware requirements.
When increasing train.num_gpus, also set train.gpu_ids to the same visible
device range, or distributed startup can be inconsistent.
AutoML runs training — all Training Requirements above apply. For no-input
local smoke runs, use DINO_AUTOML_PROFILE. For training-log-only scoring, use
val_mAP50 (extracted from Validation mAP50) or val_mAP. When each
recommendation is scored through the standalone evaluate action, use
test_mAP50 or test_mAP with direction="maximize".
See references/dino-tuning-multigpu.md for the full multi-GPU spec-consistency
rule (8-GPU example, NCCL timeout note) and the full AutoML/HPO notes (metric
selection, metric_extractor, recommended hyperparameters, weight_decay
behavior, dense-dataset resume guidance, and parent-model inference mappings).
Common failures include CUDA OOM (reduce batch_size), missing val_data_sources
(FileNotFoundError at startup — always supply val), num_classes too low (CUDA device-side assert), and the parent dino gen_trt_engine / dino convert PyT-CLI
restrictions.
See references/dino-actions-errors.md for the complete error-pattern catalog
with diagnostics and fixes.
Model-specific inference mappings belong in this MD file. For
parent_model/parent_model_folder, pass the upstream train/export/AutoML
child job id as the parent job id; list the parent result folder, filter
checkpoint artifacts, and select the resolved model.
See references/dino-tuning-multigpu.md for the full inference-mapping table
(per action: parent_model, key, output_dir, ptm_if_no_resume_model,
resume_model, create_onnx_file) and the TensorRT-mapping note. TensorRT
mappings live in the deploy workflow, not the PyT model skill.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 31 other files (references) in skills/tao-train-dino of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
Tao Train Dino next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Train Dino this skillNVIDIA/skills | 3.5k | — | ~2.8k | Automated safety check: Notes | Apache-2.0 | |
| Yolo Detection 2026SharpAI/DeepCamera | 3.1k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit | 1.1k | — | ~3.1k | Automated safety check: Pass | Custom licence | |
| Senior Computer Visionalirezarezvani/claude-skills | 28k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Senior Computer Visionborghei/Claude-Skills | 886 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Mindspeed Mm Vlmascend-ai-coding/awesome-ascend-skills | 174 | — | ~5.2k | Automated safety check: Pass | None |
SharpAI/DeepCamera
YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.
matlab/matlab-agentic-toolkit
Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.
alirezarezvani/claude-skills
Computer vision engineering skill for object detection, image segmentation, and visual AI systems.
borghei/Claude-Skills
Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment.
ascend-ai-coding/awesome-ascend-skills
Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Tao Train Dino is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection.
Tao Train Dino fits situations like: running inference for a TAO DINO detector; phrases include train DINO; DETR object detection; TAO 2D detection.
Run `npx skills add NVIDIA/skills --skill tao-train-dino -a claude-code`. Or copy the skill folder (skills/tao-train-dino in NVIDIA/skills) into .claude/skills/tao-train-dino in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-train-dino -a codex`. Or copy the skill folder (skills/tao-train-dino in NVIDIA/skills) into .agents/skills/tao-train-dino in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-dino -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-dino, .gemini/skills/tao-train-dino, .github/skills/tao-train-dino and .opencode/skills/tao-train-dino in your project.
SKILL.md names no scripts, command-line tools or credentials: Tao Train Dino is instructions for the agent only. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Train Dino is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 17k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Train Dino: Yolo Detection 2026 (SharpAI/DeepCamera, 3.1k stars), Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Senior Computer Vision (alirezarezvani/claude-skills, 28k stars) and Senior Computer Vision (borghei/Claude-Skills, 886 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.