Yolo Detection 2026
SharpAI/DeepCamera
YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.
Co-DETR (CoDINO) for object detection. An agent skill from NVIDIA/skills.
$ npx skills add NVIDIA/skills --skill tao-train-codetr -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-train-codetr --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-codetr .claude/skills/tao-train-codetr && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-train-codetr" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-codetr into .claude/skills/tao-train-codetr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-codetr", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-train-codetrType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-train-codetr -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-train-codetr --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-train-codetr .agents/skills/tao-train-codetr && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-train-codetr" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-codetr into .agents/skills/tao-train-codetr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-codetr", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-codetr -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-train-codetr --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-train-codetr .cursor/skills/tao-train-codetr && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-train-codetr" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-codetr into .cursor/skills/tao-train-codetr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-codetr", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-train-codetr--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-train-codetr -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-train-codetr --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-train-codetr .gemini/skills/tao-train-codetr && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-train-codetr" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-codetr into .gemini/skills/tao-train-codetr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-codetr", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-train-codetrInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-train-codetr -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-train-codetr .github/skills/tao-train-codetr && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-codetr" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-codetr into .github/skills/tao-train-codetr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-codetr", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-codetr -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-train-codetr --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-train-codetr .opencode/skills/tao-train-codetr && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-codetr" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-codetr into .opencode/skills/tao-train-codetr/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-codetr", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-train-codetrCo-DETR (CoDINO) for object detection. An agent skill from NVIDIA/skills.
Tao Train Codetr is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Co-DETR (CoDINO) for object detection. A DETR-family detector with collaborative hybrid assignment — auxiliary one-to-many heads supervise the encoder during training, giving strong closed-set accuracy at high inference cost. Use when training, evaluating, or running inference for a TAO Co-DETR model. Trigger phrases include "train Co-DETR", "run CoDINO", "codetr inference", "collaborative DETR", "autolabel detections with Co-DETR".
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.
It sits in AI & LLM Engineering, covering Computer vision. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockerpython3huggingface-cliFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires docker + nvidia-container-toolkit.
From compatibility in the SKILL.md frontmatter.
Tao Train Codetr loads about 4.8k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 113 tokens; SKILL.md has 1,845 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,845 words, ~4,767 tokens.
.claude/skills/tao-train-codetr/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Co-DETR trains a DETR detector alongside auxiliary one-to-many assignment heads (num_co_heads), which supervise the encoder more densely than DETR's one-to-one matching alone. The auxiliary heads exist only during training; inference runs the primary DETR head. Backbones are large by default — vit_large_codetr for training, swin_large_patch4_window7_224 for the reference inference/eval configs — so this is an accuracy-first model, not a latency-first one.
The codetr console script is not registered in the TAO PyTorch images checked so far
(7.0.1-pyt and the 2026.7.31-rc-12-multiarch nightly), even though the module itself
ships. So a bare codetr invocation fails while the network is perfectly usable.
Probe in two steps and use whichever works:
# 1. console script (preferred when present)
docker run --rm "$TAO_PYT_IMAGE" codetr --help >/dev/null 2>&1 && CODETR="codetr"
# 2. module fallback — works whenever the package is installed
[ -z "${CODETR:-}" ] && docker run --rm --entrypoint sh "$TAO_PYT_IMAGE" \
-c 'python3 -c "import nvidia_tao_pytorch.cv.codetr"' >/dev/null 2>&1 \
&& CODETR="python3 -m nvidia_tao_pytorch.cv.codetr.entrypoint.codetr"
[ -z "${CODETR:-}" ] && { echo "FATAL: Co-DETR not available in $TAO_PYT_IMAGE"; exit 1; }Both forms take identical arguments — the module entrypoint reports itself as codetr and
exposes the same {train, evaluate, inference, default_specs} subtasks. Substitute $CODETR
wherever this document writes codetr.
Only if both probes fail does this image genuinely lack Co-DETR. Then stop and ask which image to use — do not silently substitute another detector.
Generated TAO Core schemas are not yet packaged for this model, so schemas/<action>.schema.json and references/spec_template_<action>.yaml are absent except for the hand-written references/spec_template_inference.yaml. Read spec keys from the upstream experiment specs in nvidia_tao_pytorch/cv/codetr/experiment_specs/ (train.yaml, eval.yaml, inference.yaml, export.yaml) until a maintainer regenerates them. AutoML is therefore not runnable for this model.
This model is not AutoML-enabled (automl_enabled: false in references/skill_info.yaml). Run train directly; do not route through tao-skill-bank:tao-run-automl. Never add an automl_policy or workflow: key to the spec — TAO's Hydra ExperimentConfig schema rejects unknown top-level keys at config-merge time.
The packaged Co-DETR CLI exposes train, evaluate, and inference. This skill exposes all three. export and TensorRT deploy flows are not available for this model.
Every action follows the standard TAO form, where anything after the spec is a Hydra override:
$CODETR <action> -e /absolute/path/spec.yaml [key=value ...]results_dir auto-appends the action name: passing results_dir=X writes to X/train/, X/evaluate/, or X/inference/. Never append the subdirectory yourself.
| Action | Spec Key | Files |
|---|---|---|
| train | dataset.train_data_sources | image_dir, json_file (COCO) |
| train | dataset.val_data_sources | image_dir, json_file (COCO) |
| evaluate | dataset.test_data_sources | image_dir, json_file (COCO) |
| inference | dataset.infer_data_sources.image_dir | directory of images |
| inference | dataset.infer_data_sources.classmap | newline-separated class list, one name per line |
dataset.num_classes is required for every action and must match both the classmap length and the checkpoint head.
# $CODETR is either `codetr` or the module form — see ## Availability
$CODETR inference -e "$SPEC" \
inference.checkpoint=/abs/codetr.pth \
dataset.infer_data_sources.image_dir=/abs/images \
dataset.infer_data_sources.classmap=/abs/classmap.txt \
results_dir=/abs/results \
inference.conf_threshold=0.3 \
inference.num_gpus=8references/spec_template_inference.yaml is a starting point.
Two fields control where training starts:
| Field | Holds |
|---|---|
model.pretrained_backbone_path | a backbone-only checkpoint initialising model.backbone |
train.pretrained_model_path | a full Co-DETR (or compatible DETR-family) checkpoint to fine-tune from |
Prefer train.pretrained_model_path with the COCO detector below. It carries a trained
backbone and trained detection heads, so it is the better starting point for fine-tuning on a
new class set, and it is the checkpoint this skill is verified against.
A backbone-only checkpoint must match the architecture exactly. vit_large_codetr is ViT-L/16
— readable from the checkpoint as backbone.patch_embed.proj.weight (1024, 3, 16, 16), a 16x16
patch embedding. A ViT-L/14 checkpoint such as
nvidia/tao/pretrained_dinov2_classification_imagenet:vit_large_patch14_dinov2 has a
(1024, 3, 14, 14) embedding and cannot load into it; verify the patch size of any candidate
against the target backbone before using it. For the Swin, FAN, ResNet and EfficientViT
backbones, use their own matching checkpoints.
Left null, the backbone initialises randomly and needs substantially longer training.
Inference needs a full detector checkpoint, not a backbone — inference.checkpoint wants trained detection heads, which a backbone does not have. Training emits model_epoch_<N>.pth into train.results_dir.
No Co-DETR detector appears under nvidia/tao on NGC. The upstream authors publish one on HuggingFace, and it loads into this build directly:
huggingface-cli download zongzhuofan/co-detr-vit-large-coco pytorch_model.pth --local-dir <dir>zongzhuofan/co-detr-vit-large-coco — Apache-2.0, a single 2.8 GB pytorch_model.pth. Verified against 7.0.1-pyt: it is the ViT-Large COCO-80 model, and pairs with
model:
backbone: vit_large_codetr
num_queries: 1500
num_feature_levels: 5
return_interm_indices: [0, 1, 2, 3, 4]
num_co_heads: 1
dataset:
num_classes: 80Loading reports 66 missing, 0 unexpected — the missing keys are the training-only collaborative heads and are expected. Use the COCO-80 classmap with it, and fold to your own classes with inference.category_mapping (above) rather than post-processing the labels.
A ready-made COCO-80 classmap ships with this skill at references/coco80_classmap.txt — use it directly for any COCO-trained checkpoint. It is extracted from the container's own METAINFO['classes'] in nvidia_tao_pytorch/cv/deformable_detr/model/post_process.py, which is the source of truth; do not retype a COCO list from memory, since the VOC-style names (aeroplane, motorbike) and the 91-id-with-gaps ordering both look plausible and both mis-map every class silently.
dataset.infer_data_sources.classmap is a plain-text file, one class name per line, in category_id order starting at 1 (the first foreground class). Its length must equal the number of foreground classes the model predicts — 80 for a COCO-trained checkpoint with dataset.contiguous_labels: True. These names are what inference.color_map and inference.category_mapping refer to.
model.* must match the checkpoint or loading fails. Read the architecture out of the
checkpoint rather than guessing — every field below is recoverable:
sd = torch.load(ckpt, map_location="cpu", weights_only=False)["state_dict"]
# backbone.patch_embed.proj.weight (1024,3,16,16) -> ViT-L/16 -> vit_large_codetr
# query_head.transformer.query_embed.weight (1500,256) -> num_queries 1500
# max index in query_head...{encoder,decoder}.layers.N. -> enc_layers / dec_layers
# query_head...cls_branches.0.weight (80,256) -> dataset.num_classes 80
# roi_head.<i>. indices present -> num_co_heads (count)
# neck.p2..p6 -> num_feature_levels 5A correct load reports 0 unexpected:
Missing keys: 66 total (66 expected for collab heads / buffers, 0 unexpected)Missing keys are normal — the collaborative heads are training-time only. Unexpected keys are not: they mean the spec describes a different architecture than the checkpoint holds.
Confirmed against a 2.8 GB vit_large_codetr COCO checkpoint on 7.0.1-pyt:
results_dir: /abs/results
model:
backbone: vit_large_codetr
num_queries: 1500
num_feature_levels: 5
return_interm_indices: [0, 1, 2, 3, 4]
two_stage_type: standard
num_co_heads: 1
hidden_dim: 256
nheads: 8
enc_layers: 6
dec_layers: 6
dim_feedforward: 2048
dataset:
num_classes: 80
batch_size: 2
augmentation:
fixed_padding: true
fixed_random_crop: 1536 # REQUIRED by vit_large_codetr — see below
input_mean: [0.485, 0.456, 0.406]
input_std: [0.229, 0.224, 0.225]
test_random_resize: 1280
infer_data_sources:
image_dir: /abs/images
classmap: /abs/coco_classmap.txt
inference:
checkpoint: /abs/codetr_pytorch_model.pth
conf_threshold: 0.3
input_width: 640
input_height: 640
num_gpus: 2
category_mapping:
car: ["car", "bus", "truck"]Two fields in the online Co-DETR spec example are not in this build's schema and fail the
Hydra merge with Key '<name>' not in '<Config>':
| Field | Status in 7.0.1-pyt |
|---|---|
dataset.contiguous_labels | absent from DINODatasetConfig |
dataset.augmentation.pad_size_divisor | absent from DINOAugmentationConfig |
Co-DETR reuses DINO's dataset and augmentation configs, so enumerate the real fields before trusting an example:
python3 -c "
from nvidia_tao_pytorch.config.dino.dataset import DINODatasetConfig, DINOAugmentationConfig
import dataclasses
print([f.name for f in dataclasses.fields(DINODatasetConfig)])
print([f.name for f in dataclasses.fields(DINOAugmentationConfig)])"vit_large_codetr hard-requires dataset.augmentation.fixed_random_crop — the ViT backbone
needs a fixed input size, and build_nn_model.py raises without it. It is not optional at
inference despite the name suggesting a training-time augmentation.
| Parameter | Notes |
|---|---|
model.backbone | vit_large_codetr (train default) or swin_large_patch4_window7_224 (inference/eval default). Must match the checkpoint. |
model.num_queries | 1500 for the ViT config, 900 for Swin. Paired with the backbone — do not mix. |
model.num_feature_levels / return_interm_indices | 5 / [0,1,2,3,4] for ViT; 4 / [1,2,3,4] for Swin. |
model.num_co_heads | Auxiliary collaborative heads. 2 for the ViT train config, 1 for the reference inference config. Training-time only. |
model.soft_nms_enabled | Train config enables soft-NMS (linear, IoU 0.8) to match the original Co-DETR test config. |
inference.conf_threshold | Detections below this are dropped at write time. TAO default 0.5. |
dataset.num_classes | Must match classmap length and checkpoint head. |
Class indexing. The reference Co-DETR checkpoint uses 0-indexed class labels — TAO's inference path sets start_from_one=False. Verify the offset before mapping predicted class names onto another schema's integer ids.
With results_dir=X:
| Artifact | Location |
|---|---|
| KITTI detection labels | X/inference/labels/<image_stem>.txt |
| Annotated images | X/inference/images_annotated/ |
| Run config + status | X/inference/experiment.yaml, X/inference/status.json |
Label lines are KITTI-style — 15 fields plus a trailing confidence, boxes absolute xyxy:
<class_name> 0.00 0 0.00 <x1> <y1> <x2> <y2> 0.00 0.00 0.00 0.00 0.00 0.00 0.00 <score>Co-DETR predicts whatever vocabulary its checkpoint was trained on — COCO-80 for the reference checkpoint, in the order given by dataset.infer_data_sources.classmap. Emitted labels carry class names, not indices.
inference.category_mappingWhen a downstream model trains on a smaller set, fold here, not afterwards:
inference:
category_mapping:
bicycle: ["bicycle", "motorcycle"]
car: ["car", "bus", "truck"]
person: ["person"]Values are lists. Semantics, from model/category_mapping.py:
| Behaviour | Detail |
|---|---|
| unmapped originals | dropped |
| a name in two groups | keeps the first, logs a warning |
| a name absent from the classmap | warns, ignores — matching is exact, including case |
| an empty remap | raises ValueError |
| output category ids | 0..K-1 in the order the mapping is written |
Then it runs apply_category_mapping_groupnms — per-output-category soft-NMS after the merge. This is why folding here beats renaming labels afterwards: one object detected as both truck and car becomes two boxes of the same class the instant those fold together, and only a post-fold NMS removes the duplicate. A downstream rename ships the duplicates onward.
That dedup needs model.soft_nms_enabled: True, which is not the default. The schema ships it False, and with it off the fold renames classes without merging anything — the same result a downstream rename would give, duplicates included. Folding {vehicle: ["car", "bus", "truck"]} over 8 traffic images at conf_threshold: 0.05:
model.soft_nms_enabled | boxes emitted |
|---|---|
False (schema default) | 277 — exactly 146 car + 102 truck + 29 bus, nothing merged |
True | 151 — 126 duplicates removed |
Set it whenever category_mapping groups classes that the detector confuses with each other, which for COCO vehicles it reliably does. soft_nms_iou_threshold (default 0.8) controls how aggressively the merge happens.
inference.color_map keys and the class names in the emitted labels both follow the output categories once this is set.
An empty label file is meaningful and is still written: in KITTI it means "this image has no objects".
KITTI is rarely the final format. Both conversions are published in the TAO data-services container via annotations convert -e <spec>:
data.input_format: KITTI, data.output_format: COCO, plus kitti.{image_dir, label_dir, mapping}. kitti.mapping is a YAML list of single-key dicts whose values are lists — - car: [car], never - car: car. The converter builds its lookup with {label: k for k, v in cat_map.items() for label in v}, iterating the value, so a bare string yields the class names c, a, r: nothing matches, every box is dropped, and the run still prints Execution status: PASS and exits 0.data.input_format: COCO, data.output_format: ODVG, plus coco.ann_file. Emits <name>_odvg.jsonl and <name>_odvg_labelmap.json.There is no direct KITTI → ODVG conversion; chain the two.
Set train.num_gpus / inference.num_gpus. gpu_spec_key is train.num_gpus. Large backbones make this model memory-hungry — reduce dataset.batch_size before reducing GPU count when hitting OOM.
Accuracy-first with large backbones and high query counts (900–1500). Expect substantially higher memory and latency than DINO or RT-DETR at equal input size. train.activation_checkpoint trades compute for memory when training OOMs.
| Symptom | Cause | Fix |
|---|---|---|
codetr: command not found | Console script unregistered — expected on current builds; the module usually still ships | Use python3 -m nvidia_tao_pytorch.cv.codetr.entrypoint.codetr. Only if the import also fails is Co-DETR genuinely absent |
Nested inference/inference/ | Appended the action name to results_dir | Pass the parent directory only |
| Empty label files across the board | conf_threshold above the model's score range | Lower it (reference pipelines use 0.3) |
| Class names shifted by one | Reference checkpoint is 0-indexed | Check the offset when mapping to another schema |
| Size/shape mismatch loading checkpoint | model.backbone does not match the checkpoint | Align backbone, num_queries, num_feature_levels, return_interm_indices, num_co_heads |
Key 'X' not in 'DINODatasetConfig' / 'DINOAugmentationConfig' | Spec uses a documented field absent from this build | Enumerate the real dataclass fields; contiguous_labels and pad_size_divisor are documented but not in 7.0.1 |
vit_large_codetr requires dataset.augmentation.fixed_random_crop | ViT backbone needs a fixed input size | Set fixed_random_crop (1536 for the reference ViT config) — required at inference too |
| Non-zero unexpected keys at load | Spec architecture differs from the checkpoint | Re-derive model.* from the checkpoint tensors; missing keys alone are normal |
| CUDA OOM in train | Large backbone + high query count | Lower dataset.batch_size; enable train.activation_checkpoint |
| Paths not found in container | Host/container path mismatch | Mount so paths are identical; confirm images are under the mount, not just the spec |
export and TensorRT deploy are not available for Co-DETR in this build. For a deployable detector, train Co-DETR as a teacher and distil into a student that supports export — tao-train-rtdetr and tao-train-dino both expose export and gen_trt_engine.
references/checkpoint-spec-pairing.md — deriving backbone, num_queries, num_feature_levels and num_classes from a checkpoint's tensors. A mismatch fails silently: PASS, exit 0, empty labels.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 9 other files (references) in skills/tao-train-codetr of NVIDIA/skills.
Open the folder on GitHubat commit 0e0d506
Tao Train Codetr next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Train Codetr this skillNVIDIA/skills | 3.5k | — | ~4.8k | Automated safety check: Notes | Apache-2.0 | |
| Yolo Detection 2026SharpAI/DeepCamera | 3.1k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit | 1.1k | — | ~3.1k | Automated safety check: Pass | Custom licence | |
| Senior Computer Visionalirezarezvani/claude-skills | 28k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Senior Computer Visionborghei/Claude-Skills | 874 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Mindspeed Mm Vlmascend-ai-coding/awesome-ascend-skills | 174 | — | ~5.2k | Automated safety check: Pass | None |
SharpAI/DeepCamera
YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.
matlab/matlab-agentic-toolkit
Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.
alirezarezvani/claude-skills
Computer vision engineering skill for object detection, image segmentation, and visual AI systems.
borghei/Claude-Skills
Computer vision engineering for object detection, segmentation, and visual AI, covering CNN and Vision Transformer architectures and ONNX/TensorRT deployment.
ascend-ai-coding/awesome-ascend-skills
Universal VLM (vision-language understanding model) training guide for Huawei Ascend NPU using MindSpeed-MM.
theneoai/awesome-skills
Elite Computer Vision Engineer skill with expertise in deep learning for images and video (CNNs, Transformers), object detection (YOLO, DETR), segmentation, OCR, and production CV deployment…
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Co-DETR (CoDINO) for object detection. An agent skill from NVIDIA/skills. Tao Train Codetr is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Co-DETR (CoDINO) for object detection.
Tao Train Codetr fits situations like: running inference for a TAO Co-DETR model; phrases include train Co-DETR; codetr inference; collaborative DETR.
Run `npx skills add NVIDIA/skills --skill tao-train-codetr -a claude-code`. Or copy the skill folder (skills/tao-train-codetr in NVIDIA/skills) into .claude/skills/tao-train-codetr in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-train-codetr -a codex`. Or copy the skill folder (skills/tao-train-codetr in NVIDIA/skills) into .agents/skills/tao-train-codetr in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-codetr -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-codetr, .gemini/skills/tao-train-codetr, .github/skills/tao-train-codetr and .opencode/skills/tao-train-codetr in your project.
Going by SKILL.md and its folder, Tao Train Codetr needs the command-line tools its instructions call (docker, python3 and huggingface-cli). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Train Codetr is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Train Codetr: Yolo Detection 2026 (SharpAI/DeepCamera, 3.1k stars), Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars), Senior Computer Vision (alirezarezvani/claude-skills, 28k stars) and Senior Computer Vision (borghei/Claude-Skills, 874 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.