Official agent skill

Tao Train Dino

by NVIDIA in NVIDIA/skills

DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Train Dino

skills CLI
$ npx skills add NVIDIA/skills --skill tao-train-dino -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-train-dino --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-dino .claude/skills/tao-train-dino && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-train-dino
GitHub stars
3.5k
Token cost
~2.8k tokens
SKILL.md length
1,194 words
Files
32 (incl. references)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection.

  • Works in 3 steps: Train dataset URI — S3 path to… → Validation dataset URI — S3 path to… → num_classes — How many object classes?…
  • Running inference for a TAO DINO detector
  • SKILL.md covers When To Use, Reference Map, Dataclass Schemas and Train Action Policy, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Tao Train Dino is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Transformer-based detector with denoising training, multi-scale features, and optional distillation support. Use when training, evaluating, exporting, distilling, quantizing, or running inference for a TAO DINO detector. Trigger phrases include "train DINO", "DETR object detection", "TAO 2D detection", "DINO with distillation".

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 34 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.

It sits in AI & LLM Engineering, covering Computer vision. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Running inference for a TAO DINO detector
  • Phrases include train DINO
  • DETR object detection
  • TAO 2D detection

Example prompts

  • “train DINO”
  • “DETR object detection”
  • “TAO 2D detection”
  • “/tao-train-dino”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit.
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Train dataset URI — S3 path to COCO-format training data
  2. Validation dataset URI — S3 path to COCO-format val data (can be same as train)
  3. num_classes — How many object classes? Default 91 (COCO). Must be >= max(category_id) + 1. Too low causes CUDA error: device-side assert…

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Train Dino loads about 2.8k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 105 tokens; SKILL.md has 1,194 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~20k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,194 words, ~2,790 tokens.

Download SKILL.mdSave it as .claude/skills/tao-train-dino/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.
name
tao-train-dino
description
DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Transformer-based detector with denoising training, multi-scale features, and optional distillation support. Use when training, evaluating, exporting, distilling, quantizing, or running inference for a TAO DINO detector. Trigger phrases include "train DINO", "DETR object detection", "TAO 2D detection", "DINO with distillation".
allowed-tools
Read, Bash
compatibility
Requires docker + nvidia-container-toolkit.
license
Apache-2.0
metadata.tags
object, detection
metadata.version
0.1.0
metadata.author
NVIDIA Corporation

DINO

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Transformer-based detector with denoising training, multi-scale features, and optional distillation support.

Uses pretrained backbone weights (e.g. ResNet-50 ImageNet). Set model.pretrained_backbone_path for backbone-only or train.pretrained_model_path for full model.

When To Use

Train, evaluate, export, distill, quantize, or run inference for a TAO DINO 2D object detector.

For TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and TensorRT inference), read references/tao-deploy-dino.md first. Deploy spec templates live in this skill's references/ folder with the spec_template_deploy_*.yaml prefix.

Reference Map

  • references/dino-data-specs.md — dataset contracts, per-action dataset requirements, per-action spec-override examples (train, evaluate, export, deploy/gen_trt_engine, inference, quantize, distill), data-source arrays, checkpoint inference, and dataset layout.
  • references/dino-actions-errors.md — important parameters, default values, evaluate/export defaults, hardware, and the full error-pattern catalog.
  • references/dino-tuning-multigpu.md — full AutoML/HPO notes (metrics, hyperparameters, extractor) and multi-GPU spec consistency.
  • references/tao-deploy-dino.md — TensorRT deploy workflow.
  • references/detailed-guide.md — map to the detailed model guide.

Dataclass Schemas

Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML still requires schemas/train.schema.json and references/spec_template_train.yaml to exist and parse. Use the packaged train schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.

Train Action Policy

This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.

Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.

Training Requirements

The agent MUST read this section before generating any training or AutoML script for DINO.

  • Dataset type: object_detection
  • Formats: coco, coco_raw
  • Accepted dataset intents: training, evaluation, testing, calibration
  • AutoML metric contract: for the default evaluation-backed workflow, use test_mAP50 with maximize direction. Use test_mAP only when the user explicitly requests COCO/paper-style mAP.
  • Training monitoring metrics: val_mAP50 (logged as Validation mAP50) for quick operational checks; val_mAP for COCO/paper-style benchmark comparisons.
  • Evaluate action metrics: test_mAP50 for AP50 and test_mAP for COCO mAP. AutoML workflows that score recommendations with the standalone evaluate action must use the corresponding test_* KPI.

Required datasets — MUST resolve both:

DatasetRequiredWhy
Train dataset URIYesTraining data (COCO format)
Validation dataset URIYes — ALWAYSDINO unconditionally builds a val dataloader. Omitting val_data_sources causes FileNotFoundError at startup regardless of the metric or workflow. If the user has no separate eval split, reuse the train URI.

Required inputs before generating any training spec:

  1. Train dataset URI — S3 path to COCO-format training data
  2. Validation dataset URI — S3 path to COCO-format val data (can be same as train)
  3. num_classes — How many object classes? Default 91 (COCO). Must be >= max(category_id) + 1. Too low causes CUDA error: device-side assert triggered.

Resolve these from the user request or the default profile below. Prompt only for values that are still missing after applying the profile rules.

Bankable local default profile for DINO AutoML smoke runs:

Use this profile only when the user asks to run DINO AutoML and does not provide dataset or class-count inputs. This profile is intentionally small and local to this skill bank; it is for smoke/iteration runs, not a production benchmark. Do not search previous runners, logs, session state, shell history, or the home directory to recover these values.

python
DINO_AUTOML_PROFILE = {
    "train_dataset_uri": "s3://nvcf-storage-handling/data/tao_od_synthetic_subset_train_no_convert",
    "validation_dataset_uri": "s3://nvcf-storage-handling/data/tao_od_synthetic_subset_val_no_convert",
    "object_classes": 4,
    "dataset_num_classes": 5,
    "image_archive": "images.tar.gz",
    "annotation_file": "annotations.json",
    "max_recommendations": 10,
    "train_num_epochs": 10,
    "train_checkpoint_interval": 10,
    "train_validation_interval": 1,
    "train_num_gpus": 1,
}

If the user supplies any dataset URI or class-count value, prefer the user value and ask for any remaining required DINO value. Do not partially mix a user's custom dataset with this profile's class count unless the user confirms it.

Do not prompt for image layout for the standard DINO dataset. The standard TAO DINO dataset artifact is images.tar.gz plus annotations.json. Use images.tar.gz in the remote image_dir spec override. The SDK downloads the archive and rewrites the runtime spec to the extracted folder named after the archive stem (images.tar.gz -> images). Only deviate if the user explicitly provides a different image artifact name.

Show full SKILL.md (384 more words)Show less

Core Workflow

DINO supports train, evaluate, export, distill, quantize, and inference. Data-source overrides are mandatory for every action — DINO's config.json has empty data_sources because the runner cannot auto-resolve array-of-objects spec keys. The agent MUST construct data source paths and include them in spec_overrides.

See references/dino-data-specs.md for the per-action dataset requirements table, the standard dataset artifact (images.tar.gz + annotations.json) and runtime folder rewrite rules, and the complete per-action spec_overrides examples for train, evaluate, export, deploy/gen_trt_engine, inference, quantize, and distill — including checkpoint inference via parent_model, the results_dir/train/ checkpoint location, and the distillation FAN-teacher / student rules.

Important Parameters And Defaults

Default values: num_epochs=10, batch_size=4, learning_rate=2e-4, lr_backbone=2e-5, num_classes=91, backbone=resnet_50.

  • dataset.num_classes: Default 91 (COCO). Must be >= max(category_id) + 1. Too low causes CUDA error: device-side assert triggered. Set as <num_classes> + 1 in spec overrides.
  • num_epochs: default 10 (quick iteration); real datasets typically need 30-50+ epochs for good mAP.

See references/dino-actions-errors.md for the full parameter list (backbone options, train.optim.lr/lr_steps, model.num_queries, batch_size), default values, evaluate defaults, export defaults (input 960x544, opset 17, TRT data types, workspace 1024 MB), and hardware requirements.

Multi-GPU And AutoML / HPO

When increasing train.num_gpus, also set train.gpu_ids to the same visible device range, or distributed startup can be inconsistent.

AutoML runs training — all Training Requirements above apply. For no-input local smoke runs, use DINO_AUTOML_PROFILE. For training-log-only scoring, use val_mAP50 (extracted from Validation mAP50) or val_mAP. When each recommendation is scored through the standalone evaluate action, use test_mAP50 or test_mAP with direction="maximize".

See references/dino-tuning-multigpu.md for the full multi-GPU spec-consistency rule (8-GPU example, NCCL timeout note) and the full AutoML/HPO notes (metric selection, metric_extractor, recommended hyperparameters, weight_decay behavior, dense-dataset resume guidance, and parent-model inference mappings).

Error Patterns

Common failures include CUDA OOM (reduce batch_size), missing val_data_sources (FileNotFoundError at startup — always supply val), num_classes too low (CUDA device-side assert), and the parent dino gen_trt_engine / dino convert PyT-CLI restrictions.

See references/dino-actions-errors.md for the complete error-pattern catalog with diagnostics and fixes.

Spec Param / Parent Model Inference

Model-specific inference mappings belong in this MD file. For parent_model/parent_model_folder, pass the upstream train/export/AutoML child job id as the parent job id; list the parent result folder, filter checkpoint artifacts, and select the resolved model.

See references/dino-tuning-multigpu.md for the full inference-mapping table (per action: parent_model, key, output_dir, ptm_if_no_resume_model, resume_model, create_onnx_file) and the TensorRT-mapping note. TensorRT mappings live in the deploy workflow, not the PyT model skill.

Deployment

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 31 other files (references) in skills/tao-train-dino of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/detailed-guide.md
  • references/dino-actions-errors.md
  • references/dino-data-specs.md
  • references/dino-tuning-multigpu.md
  • references/skill_info.yaml
  • references/spec_template_deploy_evaluate.yaml
  • references/spec_template_deploy_gen_trt_engine.yaml
  • references/spec_template_deploy_inference.yaml
  • references/spec_template_distill.yaml
  • references/spec_template_evaluate.yaml
  • references/spec_template_export.yaml
  • references/spec_template_gen_trt_engine.yaml
  • references/spec_template_inference.yaml
  • references/spec_template_quantize.yaml
  • … and 14 more

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Tao Train Dino next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Train Dino compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Train Dino this skillNVIDIA/skills3.5k—~2.8kAutomated safety check: NotesApache-2.0
Yolo Detection 2026SharpAI/DeepCamera3.1k—~1.5kAutomated safety check: PassMIT
Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassCustom licence
Senior Computer Visionalirezarezvani/claude-skills28k1 repos~3.2kAutomated safety check: PassMIT
Senior Computer Visionborghei/Claude-Skills886—~1.8kAutomated safety check: PassMIT
Mindspeed Mm Vlmascend-ai-coding/awesome-ascend-skills174—~5.2kAutomated safety check: PassNone

Similar skills

  • Yolo Detection 2026

    SharpAI/DeepCamera

    YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.

    3.1k GitHub stars~1.5k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Matlab Use Visual Inspection

    matlab/matlab-agentic-toolkit

    Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

    1.1k GitHub stars~3.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Senior Computer Vision

    alirezarezvani/claude-skills

    Computer vision engineering skill for object detection, image segmentation, and visual AI systems.

    28k GitHub starsUsed in 1 repo~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Senior Computer Vision

    borghei/Claude-Skills

    Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment.

    886 GitHub stars~1.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Mindspeed Mm Vlm

    ascend-ai-coding/awesome-ascend-skills

    Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM.

    174 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Train Dino

What does Tao Train Dino do?

DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection. Tao Train Dino is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. DINO (DETR with Improved DeNoising Anchor Boxes) for 2D object detection.

When should I use Tao Train Dino?

Tao Train Dino fits situations like: running inference for a TAO DINO detector; phrases include train DINO; DETR object detection; TAO 2D detection.

How do I install Tao Train Dino in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-train-dino -a claude-code`. Or copy the skill folder (skills/tao-train-dino in NVIDIA/skills) into .claude/skills/tao-train-dino in your project. Claude Code loads it when a task matches its description.

How do I install Tao Train Dino in Codex?

Run `npx skills add NVIDIA/skills --skill tao-train-dino -a codex`. Or copy the skill folder (skills/tao-train-dino in NVIDIA/skills) into .agents/skills/tao-train-dino in your project. Codex loads it when a task matches its description.

Can I use Tao Train Dino in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-dino -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-dino, .gemini/skills/tao-train-dino, .github/skills/tao-train-dino and .opencode/skills/tao-train-dino in your project.

What does Tao Train Dino need to run?

SKILL.md names no scripts, command-line tools or credentials: Tao Train Dino is instructions for the agent only. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..

Does Tao Train Dino access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tao Train Dino safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Train Dino use?

Tao Train Dino is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Train Dino use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 17k tokens, read only when the agent opens those files.

What are the alternatives to Tao Train Dino?

Skills that share tags, products or a category with Tao Train Dino: Yolo Detection 2026 (SharpAI/DeepCamera, 3.1k stars), Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Senior Computer Vision (alirezarezvani/claude-skills, 28k stars) and Senior Computer Vision (borghei/Claude-Skills, 886 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Train Dino?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.