Official agent skill

Tao Data Io

by NVIDIA in NVIDIA/skills

The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective +…

OfficialApache-2.0Auto-check: warningsDevOps & Cloud

Install Tao Data Io

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add NVIDIA/skills --skill tao-data-io -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-data-io --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-data-io .claude/skills/tao-data-io && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-data-io
GitHub stars
3.5k
Used in
1 other repo
Token cost
~1.5k tokens
SKILL.md length
590 words
Files
8 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective +…

  • Phrases include stage inputs
  • SKILL.md covers Credentials (env vars; values…, Storage strategy (agent…, Verify-before-launch gate (the… and Input staging, plus 3 more sections
  • Runs Python scripts from its folder; calls aws, python and kubectl; needs AWS_ACCESS_KEY_ID and AWS_SECRET_ACCESS_KEY
  • Mount the dataset

What it does

Tao Data Io is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective + archive extract + HF/NGC PTM), maps credentials to env, routes outputs 3-way with upload-excludes, and runs the compute-frame verify gate. A support skill other platform skills (docker, kubernetes, slurm, brev, virtualenv) call to get data to and from the compute container without the TAO SDK. Trigger phrases include "stage…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires aws CLI or s5cmd on the staging host, plus Python 3.10+ with boto3 and pandas/pyarrow for annotation-selective download. No nvidia-tao-sdk, no…

It sits in DevOps & Cloud, covering File uploads and storage, Container orchestration and Containers. It works with NVIDIA AI Platform, Docker, Kubernetes and Amazon Web Services. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Phrases include stage inputs
  • Mount the dataset
  • Upload TAO results
  • Download only referenced files

Example prompts

  • “stage inputs”
  • “mount the dataset”
  • “upload TAO results”
  • “/tao-data-io”

Requirements

  • Python 3
  • Docker
  • A credential in AWS_SECRET_ACCESS_KEY
  • A credential in ACCESS_KEY
  • Compatibility (from SKILL.md): Requires aws CLI or s5cmd on the staging host, plus Python 3.10+ with boto3 and pandas/pyarrow for annotation-selective download. No nvidia-tao-sdk, no fsspec/s3fs. Credentials are read from the process environment, whether exported in the user's shell or sourced from a user-approved env file.
  • Pre-approved tools (allowed-tools): Read, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • aws
    • python
    • kubectl
    • huggingface-cli

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use aws and kubectl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AWS_ACCESS_KEY_ID
    • AWS_SECRET_ACCESS_KEY
    • ACCESS_KEY
    • SECRET_KEY
    • HF_TOKEN
    • NGC_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires aws CLI or s5cmd on the staging host, plus Python 3.10+ with boto3 and pandas/pyarrow for annotation-selective download. No nvidia-tao-sdk, no fsspec/s3fs. Credentials are read from the process environment, whether exported in the user's shell or sourced from a user-approved env file.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Data Io loads about 1.5k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 170 tokens; SKILL.md has 590 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~170
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a credentials file (SSH keys, cloud or package-manager tokens)SKILL.md:39
    and **never** write `~/.aws/credentials`:
  • NoteMentions a .env fileSKILL.md:42
    set -a; source /path/to/.env; set +a   # omit if already exported
  • NoteMentions a .env fileSKILL.md:98
    set -a; source /path/to/.env; set +a   # omit if already exported
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 590 words, ~1,501 tokens.

Download SKILL.mdSave it as .claude/skills/tao-data-io/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
tao-data-io
description
The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective + archive extract + HF/NGC PTM), maps credentials to env, routes outputs 3-way with upload-excludes, and runs the compute-frame verify gate. A support skill other platform skills (docker, kubernetes, slurm, brev, virtualenv) call to get data to and from the compute container without the TAO SDK. Trigger phrases include "stage inputs", "mount the dataset", "upload TAO results", "download only referenced files", "resolve results_dir", "verify the container can read the data".
allowed-tools
Read, Bash
compatibility
Requires aws CLI or s5cmd on the staging host, plus Python 3.10+ with boto3 and pandas/pyarrow for annotation-selective download. No nvidia-tao-sdk, no fsspec/s3fs. Credentials are read from the process environment, whether exported in the user's shell or sourced from a user-approved env file.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
platform, storage

tao-data-io

Get data to and from the compute container. Decide the storage tier first — under strategy A (pre-positioned mount) no bytes move at all — and when a fetch is needed, move it host-side with aws/s5cmd/boto3/huggingface-cli/ngc directly — no nvidia-tao-sdk, no in-container runtime. Other platform skills call this skill to stage inputs before launch and sync outputs after. It never launches a container itself. The chosen tier is stamped into the job-record at submit.

When NOT to invoke this skill: if the inputs are already readable from the compute frame (a local path on the execution host, an existing Lustre/PVC/bind mount), that IS tier A — record it and skip this skill entirely; there is nothing to move. Air-gapped hosts: tier A is the only tier — never attempt an S3/HF/NGC fetch; anything missing (datasets, checkpoints, and the container images themselves) must be pre-positioned by the operator, and the preflight's readability check is the only data step that runs.

Credentials (env vars; values never written to disk by this skill)

S3 credentials use the officially documented AWS env vars, read from the session environment: AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, and (for S3-compatible stores) AWS_ENDPOINT_URL, AWS_DEFAULT_REGION. The aws CLI and boto3 pick the variables up natively — never run aws configure and never write ~/.aws/credentials:

bash
set -a; source /path/to/.env; set +a   # omit if already exported
aws s3 ls "s3://$S3_BUCKET_NAME/..."   # reads AWS_* from the environment

If a session provides only the legacy TAO names (ACCESS_KEY, SECRET_KEY, S3_ENDPOINT_URL, CLOUD_REGION), map them once, scoped to the command: AWS_ACCESS_KEY_ID="$ACCESS_KEY" AWS_SECRET_ACCESS_KEY="$SECRET_KEY" aws s3 ...

HF_TOKEN / NGC_KEY pass through unchanged for PTM pulls. Never pass a credential as a CLI argument (-p, --token, -e KEY=value); use --password-stdin or -e VAR (no value).

Storage strategy (agent decides; the gate verifies)

Pick per backend from what the cluster/daemon actually offers:

  • A — pre-positioned mount (PVC / NFS / Lustre / bind mount already holding the data): mount it, author the mount paths into the spec. No S3 fetch. Also the air-gap answer.
  • B — volume populated from S3: an initContainer / stage step fills a durable volume; compute mounts it; a final step drains it back to S3 if needed.
  • C — ephemeral + in-compute fetch: no persistent mount — fetch into the compute container and upload out at the end (today's K8s default).
Show full SKILL.md (242 more words)Show less

Verify-before-launch gate (the one invariant)

The path the spec references is readable from the compute frame (not the launcher's), and the output destination persists after the container exits.

Probe in the compute's frame of reference (in-container aws s3 cp/touch on the resolved results_dir, or a kubectl run/srun probe) — a green aws s3 ls on the launcher is not proof the pod can read the data (managed backends inject different creds into the compute container).

Input staging

  • Bulk folder: s5cmd cp 's3://.../*' <stage> or aws s3 sync.
  • Single file: aws s3 cp.
  • Annotation-selective (download only files referenced by an annotation): use references/selective_download.py (below).
  • Archive: tar -xzf X -C <dir> --strip-components=1 guarded by a .extracted marker (idempotent).
  • PTM (ngc:// / hf://): huggingface-cli download / ngc registry model download-version, then author the local path.

After staging, author the spec with local paths and run the verify gate.

Output routing (3-way) + upload

  • TAO_RESULTS_ROOT set → write to that mount, no upload.
  • else S3_BUCKET_NAME set → upload to s3://$S3_BUCKET_NAME/results/$TAO_JOB_ID/.
  • else → loud ephemeral warning.

SLURM: never set S3_BUCKET_NAME (Lustre-only); run any upload on the login node, not inside the GPU allocation. Upload with excludes: aws s3 sync <local>/ s3://... --exclude '.tao/*' <upload_excludes...>.

Quick Start — annotation-selective staging

Download only the files an annotation references (e.g. the video column), preserving relative paths, into a local staging dir:

bash
set -a; source /path/to/.env; set +a   # omit if already exported
python references/selective_download.py \
  --annotation /path/to/annotation.parquet \
  --key video \
  --bucket "$S3_BUCKET_NAME" --src-prefix datasets/clips \
  --dest /data/stage/clips

--key is repeatable or comma-separated; --format overrides extension inference (parquet/jsonl/json/csv).

Helpers

  • references/selective_download.py — annotation-driven selective download (boto3 + pandas). Unit tests live in references/tests/; run with python -m pytest.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (references) in skills/tao-data-io of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/selective_download.py
  • references/tests/test_selective_download.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Tao Data Io next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Data Io compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Data Io this skillNVIDIA/skills3.5k1 repos~1.5kAutomated safety check: WarnApache-2.0
Logfire Infrastructurepydantic/skills140—~1.8kAutomated safety check: PassMIT
Ksaildevantler-tech/ksail165—~1.1kAutomated safety check: PassCustom licence
Aspire DeploymentCommunityToolkit/Aspire629—~4.5kAutomated safety check: NotesMIT
Container Orchestrationaiskillstore/marketplace430—~1.4kAutomated safety check: NotesMIT
Deployment Automationaiskillstore/marketplace4301 repos~3kAutomated safety check: NotesNone

Similar skills

  • Logfire Infrastructure

    pydantic/skills

    Official

    Monitor hosts, Docker containers, Kubernetes clusters, database/queue/cache servers, and cloud-provider metrics with Pydantic Logfire — no application code required.

    140 GitHub stars~1.8k tokensUpdated 6 days ago
    DevOps & CloudAuto-check passed
  • Ksail

    devantler-tech/ksail

    Use the ksail CLI to spin up and manage Kubernetes clusters (Kind/K3d/Talos/vCluster/KWOK — local via Docker; EKS — cloud via AWS) and GitOps workloads declaratively.

    165 GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Aspire Deployment

    CommunityToolkit/Aspire

    WORKFLOW SKILL — Deploy Aspire apps from AppHost models to Docker Compose, Kubernetes, Azure, AWS, or preview Radius.

    629 GitHub stars~4.5k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Container Orchestration

    aiskillstore/marketplace

    Docker, Kubernetes, and AWS ECS/Fargate patterns. An agent skill from aiskillstore/marketplace.

    430 GitHub stars~1.4k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Deployment Automation

    aiskillstore/marketplace

    Automate application deployment to cloud platforms and servers.

    430 GitHub starsUsed in 1 repo~3k tokens
    DevOps & CloudAuto-check: notes
  • Agent Bom Scan Infra

    LeoYeAI/openclaw-master-skills

    Scan infrastructure-as-code, cloud configurations, and find secrets.

    2.2k GitHub stars~1.5k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Categories

Questions about Tao Data Io

What does Tao Data Io do?

The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective +…. Tao Data Io is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. The data-mover for TAO jobs — decides the storage tier (A pre-positioned mount with zero fetch / B volume-from-S3 / C ephemeral in-compute fetch), stages inputs (bulk + annotation-selective + archive extract + HF/NGC PTM), maps credentials to env, routes outputs 3-way with upload-excludes, and runs the compute-frame verify gate.

When should I use Tao Data Io?

Tao Data Io fits situations like: phrases include stage inputs; mount the dataset; upload TAO results; download only referenced files.

How do I install Tao Data Io in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-data-io -a claude-code`. Or copy the skill folder (skills/tao-data-io in NVIDIA/skills) into .claude/skills/tao-data-io in your project. Claude Code loads it when a task matches its description.

How do I install Tao Data Io in Codex?

Run `npx skills add NVIDIA/skills --skill tao-data-io -a codex`. Or copy the skill folder (skills/tao-data-io in NVIDIA/skills) into .agents/skills/tao-data-io in your project. Codex loads it when a task matches its description.

Can I use Tao Data Io in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-data-io -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-data-io, .gemini/skills/tao-data-io, .github/skills/tao-data-io and .opencode/skills/tao-data-io in your project.

What does Tao Data Io need to run?

Going by SKILL.md and its folder, Tao Data Io needs Python for the scripts in its folder, the command-line tools its instructions call (aws, python, kubectl and huggingface-cli) and credentials named AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, ACCESS_KEY and SECRET_KEY. Our summary lists: Python 3; Docker; A credential in AWS_SECRET_ACCESS_KEY; A credential in ACCESS_KEY. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires aws CLI or s5cmd on the staging host, plus Python 3.10+ with boto3 and pandas/pyarrow for annotation-selective download. No nvidia-tao-sdk, no fsspec/s3fs. Credentials are read from the process environment, whether exported in the user's shell or sourced from a user-approved env file..

Does Tao Data Io access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tao Data Io safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a credentials file (ssh keys, cloud or package-manager tokens). Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Tao Data Io use?

Tao Data Io is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Data Io use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.

What are the alternatives to Tao Data Io?

Skills that share tags, products or a category with Tao Data Io: Logfire Infrastructure (pydantic/skills, 140 stars), Ksail (devantler-tech/ksail, 165 stars), Aspire Deployment (CommunityToolkit/Aspire, 629 stars) and Container Orchestration (aiskillstore/marketplace, 430 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Data Io?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.