Official agent skill

Tao Validate Dataset Format

by NVIDIA in NVIDIA/skills

Run tao-daft validate to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors.

OfficialApache-2.0Auto-check: notes

Install Tao Validate Dataset Format

skills CLI
$ npx skills add NVIDIA/skills --skill tao-validate-dataset-format -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-validate-dataset-format --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-validate-dataset-format .claude/skills/tao-validate-dataset-format && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-validate-dataset-format
GitHub stars
3.5k
Token cost
~1.2k tokens
SKILL.md length
545 words
Files
6
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run tao-daft validate to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors.

  • Works in 3 steps: Format is a positional subcommand, not… → Target is --path PATH, not positional.… → Flags are per-format; run the leaf help,…
  • Non-DAFT formats
  • SKILL.md covers Quick start, Preflight, Quick Start and Purpose, plus 4 more sections
  • Calls pip and python

What it does

Tao Validate Dataset Format is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run tao-daft validate to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors. Do not use for non-DAFT formats. Use when the user asks to validate a DAFT dataset, check DAFT schema, validate a TAO dataset format, or run tao-daft validate.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires Python 3.10+ and the nvidia-tao-daft package (pip install nvidia-tao-daft).

It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Non-DAFT formats
  • The user asks to validate a DAFT dataset
  • Check DAFT schema
  • Validate a TAO dataset format

Example prompts

  • “/tao-validate-dataset-format”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.10+ and the nvidia-tao-daft package (pip install nvidia-tao-daft).
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Format is a positional subcommand, not --format
  2. Target is --path PATH, not positional. It accepts a single
  3. Flags are per-format; run the leaf help, e.g.

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.10+ and the nvidia-tao-daft package (pip install nvidia-tao-daft).

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Validate Dataset Format loads about 1.2k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 545 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 545 words, ~1,243 tokens.

Download SKILL.mdSave it as .claude/skills/tao-validate-dataset-format/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
tao-validate-dataset-format
description
Run `tao-daft validate` to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors. Do not use for non-DAFT formats. Use when the user asks to validate a DAFT dataset, check DAFT schema, validate a TAO dataset format, or run `tao-daft validate`.
allowed-tools
Read, Bash
compatibility
Requires Python 3.10+ and the nvidia-tao-daft package (pip install nvidia-tao-daft).
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
tao-daft, dataset, validation, schema

Validate a TAO DAFT Dataset

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Quick start

bash
tao-daft validate <format> --path <dataset-or-parent-dir>

<format> is a positional subcommand (e.g. metropolis-v3.0, cosmos-reason-v1.0); --path is required. Discover supported formats and per-format flags via tao-daft validate --help and the leaf --help (see "CLI conventions" below).

Preflight

bash
python -c "import nvidia_tao_daft" 2>/dev/null || {
  echo "MISSING: tao-daft not installed. Run:"
  echo "  pip install nvidia-tao-daft"
  exit 1
}

Quick Start

Discover the installed validator formats before choosing a format slug, then run validation with the target passed through --path:

bash
tao-daft --version
tao-daft validate --help
tao-daft validate <format> --help
tao-daft validate <format> --path /path/to/daft-dataset

Purpose

Drive tao-daft validate against a DAFT dataset (or a tree of them). The CLI is the spec; the skill picks subcommand + flags and explains the result.

Trigger when the user mentions "TAO DAFT", "DAFT format", validating a DAFT dataset, schema/cross-reference errors, or tao-daft validate. Do not trigger for non-DAFT layouts (COCO, YOLO, Data Factory JSONL), or for tao-daft info / tao-daft convert — those have their own skills.

If the user's opening is ambiguous, run a few --help commands first to ground yourself, then come back and confirm the task.

Prerequisites

  • nvidia-tao-daft installed (pip install nvidia-tao-daft; the wheel is enough, no source repo). Confirm with tao-daft --version.
  • A DAFT dataset, or a parent directory of them, on local disk.

Instructions

CLI conventions

tao-daft is nested argparse subcommands. Names and flags drift across versions, so discover the current surface from --help rather than trusting any list in this doc.

  1. Format is a positional subcommand, not --format: tao-daft validate <format> [flags]. List current formats via tao-daft validate --help; slugs look like metropolis-v3.0, cosmos-reason-v1.0.
  2. Target is --path PATH, not positional. It accepts a single dataset/scene or a parent directory — the validator walks the tree.
  3. Flags are per-format; run the leaf help, e.g. tao-daft validate metropolis-v3.0 --help, before choosing them. Don't assume a flag from one format exists on another.

So the loop is: tao-daft --version → tao-daft validate --help → pick format (infer if unspecified, see below) → tao-daft validate <format> --help → run → interpret.

Show full SKILL.md (223 more words)Show less
Format inference

Use directory markers, not filenames:

  • meta.json next to media/ and text/ ⇒ cosmos-reason-v1.0.
  • A directory (or nested directories) containing contextual/, typically alongside raw/ and task/ ⇒ metropolis-v3.0.
  • Neither marker present ⇒ ask the user; do not guess.
Reading errors

The CLI ends every run with a VALIDATION RESULTS block, then ✅ VALIDATION PASSED or ❌ VALIDATION FAILED, and exits non-zero on failure (safe to chain in scripts).

Output can be large on big trees — capture the full output to a file and read it in slices rather than scrolling inline.

Limitations

  • Validates DAFT only. Non-DAFT layouts (COCO, YOLO, Data Factory JSONL, etc.) belong in the upstream converter skills.
  • Supported formats are whatever tao-daft validate --help reports for the installed version; older slugs may have been retired.
  • Covers validate only. Defer to the dedicated skills for tao-daft info and tao-daft convert.
  • Don't reimplement validation in Python; the CLI is the spec.

Troubleshooting

  • tao-daft: command not found — wheel not installed in the active env. pip install nvidia-tao-daft; verify tao-daft --version.
  • error: argument --path is required — path passed positionally. Move it behind --path.
  • invalid choice: '<format>' — slug isn't wired up in this version. Re-run tao-daft validate --help and pick from the list.
  • Auto-detection (raw type / contextual set) is wrong — override via the format's scope-restriction flag; discover the name from the leaf --help.
  • CI wants warnings to fail — add --strict.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/tao-validate-dataset-format of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Tao Validate Dataset Format next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Validate Dataset Format compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Validate Dataset Format this skillNVIDIA/skills3.5k—~1.2kAutomated safety check: NotesApache-2.0
Skill InspectorNVIDIA/SkillSpector20k1 repos~1.8kAutomated safety check: PassApache-2.0
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
NEAR AI Cloud Private Inferenceinternet-court/internet-court-skill6.4k2 repos~1.3kAutomated safety check: PassCustom licence
Nemoclaw Maintainer Normalize Title TagsNVIDIA/NemoClaw23k—~693Automated safety check: PassApache-2.0

Similar skills

  • Skill Inspector

    NVIDIA/SkillSpector

    Official

    Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.

    20k GitHub starsUsed in 1 repo~1.8k tokens
    SecurityAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • NEAR AI Cloud Private Inference

    internet-court/internet-court-skill

    Shows how to call NEAR AI Cloud through an OpenAI-compatible API and verify that inference ran in a TEE, using attestation checks and signed chat responses.

    6.4k GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Remove bracketed NemoClaw tags from GitHub issue and PR titles.

    23k GitHub stars~693 tokensUpdated today
    Marketing & SEOAuto-check passed
  • Audit and implement a NemoClaw dependency version upgrade, including Hermes and base images.

    23k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Validate Dataset Format

What does Tao Validate Dataset Format do?

Run tao-daft validate to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors. Tao Validate Dataset Format is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run tao-daft validate to check NVIDIA TAO DAFT datasets for structure, schema, and cross-reference errors.

When should I use Tao Validate Dataset Format?

Tao Validate Dataset Format fits situations like: non-DAFT formats; the user asks to validate a DAFT dataset; check DAFT schema; validate a TAO dataset format.

How do I install Tao Validate Dataset Format in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-validate-dataset-format -a claude-code`. Or copy the skill folder (skills/tao-validate-dataset-format in NVIDIA/skills) into .claude/skills/tao-validate-dataset-format in your project. Claude Code loads it when a task matches its description.

How do I install Tao Validate Dataset Format in Codex?

Run `npx skills add NVIDIA/skills --skill tao-validate-dataset-format -a codex`. Or copy the skill folder (skills/tao-validate-dataset-format in NVIDIA/skills) into .agents/skills/tao-validate-dataset-format in your project. Codex loads it when a task matches its description.

Can I use Tao Validate Dataset Format in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-validate-dataset-format -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-validate-dataset-format, .gemini/skills/tao-validate-dataset-format, .github/skills/tao-validate-dataset-format and .opencode/skills/tao-validate-dataset-format in your project.

What does Tao Validate Dataset Format need to run?

Going by SKILL.md and its folder, Tao Validate Dataset Format needs the command-line tools its instructions call (pip and python). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires Python 3.10+ and the nvidia-tao-daft package (pip install nvidia-tao-daft)..

Does Tao Validate Dataset Format access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tao Validate Dataset Format safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Validate Dataset Format use?

Tao Validate Dataset Format is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Validate Dataset Format use?

About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Tao Validate Dataset Format?

Skills that share tags, products or a category with Tao Validate Dataset Format: Skill Inspector (NVIDIA/SkillSpector, 20k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Embeddings via 9Router (decolua/9router, 30k stars) and NEAR AI Cloud Private Inference (internet-court/internet-court-skill, 6.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Validate Dataset Format?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.