Official agent skill

TAO Image Embeddings

by NVIDIA in NVIDIA/skills

Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install TAO Image Embeddings

skills CLI
$ npx skills add NVIDIA/skills --skill tao-generate-image-embeddings -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-generate-image-embeddings --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-generate-image-embeddings .claude/skills/tao-generate-image-embeddings && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-generate-image-embeddings
GitHub stars
3.5k
Token cost
~2k tokens
SKILL.md length
742 words
Files
9 (incl. scripts, references, assets)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining.

  • Works in 4 steps: Verify Docker and GPU access → Resolve and pull the data-services image… → Validate the spec → …
  • Embedding a folder of images before nearest-neighbor mining
  • SKILL.md covers Inputs, Encoder Consistency, Quick Start and Generate A Spec, plus 3 more sections
  • Runs Python scripts from its folder; calls docker and python3

What it does

The skill takes a spec with an input parquet, an output parquet, a model of CLIP or SigLIP and a matching model path, which can be a Hugging Face model ID, a local snapshot directory or a TAO checkpoint, and runs the embedding command inside the TAO Data Services container. The input parquet needs a filepath column, and other columns such as labels are carried into the output. A default template and a spec validator that rejects mismatched model and path ship with it.

It stresses encoder consistency: every parquet compared in a mining run must be embedded with the same model and path, because dimensions differ (768 for the SigLIP default, 512 for CLIP ViT-B/32), the output records nothing about the encoder and mixed encoders are the usual cause of unrelated-looking results. Saving the spec beside the output is recommended. Batch size defaults to 64 and can be lowered if the GPU runs out of memory. The output feeds the tao-mine-od-images and tao-mine-nearest-neighbors skills.

When your agent uses it

  • Embedding a folder of images before nearest-neighbor mining
  • Generating SigLIP or CLIP embeddings for a dataset listed in a parquet file
  • Checking that a model and model path pair match before launching a run
  • Producing embeddings from a TAO checkpoint

Example prompts

  • “Embed the images listed in data/images.parquet with SigLIP and write data/embeddings.parquet.”
  • “Generate CLIP embeddings for my training set and keep the label column.”
  • “The GPU ran out of memory on the embedding run, so lower the batch size and retry.”
  • “Check my image_embeddings.yaml for a model and model_path mismatch.”

Requirements

  • Docker with the NVIDIA container toolkit
  • One or more CUDA GPUs
  • The TAO data-services container pinned in versions.yaml
  • Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Verify Docker and GPU access
  2. Resolve and pull the data-services image if needed
  3. Validate the spec
  4. Confirm RUN_ROOT contains the spec, the input parquet, the image files its filepath column points at, and the output directory. Mount…

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • docker
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.

    From compatibility in the SKILL.md frontmatter.

Context cost

TAO Image Embeddings loads about 2k tokens when it runs, and up to ~2.1k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 742 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 742 words, ~1,972 tokens.

Download SKILL.mdSave it as .claude/skills/tao-generate-image-embeddings/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
tao-generate-image-embeddings
description
Run TAO Data Services image embedding to turn a parquet of image filepaths into an embedding parquet using CLIP, SigLIP, or a TAO checkpoint. Use when a workflow needs embeddings before nearest-neighbor or unique-neighbor mining, or when the user asks to "embed images", "compute image embeddings", or "generate SigLIP embeddings".
allowed-tools
Read, Bash
compatibility
Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
tao, data, embeddings, mining, siglip, clip

TAO Generate Image Embeddings

Use this skill to run TAO Data Services image embedding. The skill consumes a parquet of image filepaths and writes a parquet with an embedding column. Downstream mining skills (tao-mine-od-images, tao-mine-nearest-neighbors) consume its output.

The container entrypoint is:

bash
embedding image_embeddings -e /absolute/path/to/image_embeddings.yaml

Inputs

The user can provide either an existing spec or the fields needed to generate one.

Required spec fields:

FieldMeaning
input_parquetAbsolute path to a parquet containing image filepaths.
output_parquetAbsolute path where the embedding parquet is written.
modelCLIP or SigLIP.
model_pathHuggingFace model id, local HF snapshot directory, or a TAO .pth/.ckpt checkpoint. Must match model. SigLIP: google/siglip-base-patch16-224 (768-dim, the template default). CLIP: openai/clip-vit-base-patch32 (512-dim). The validator rejects a recognizable model/path mismatch before launch.

Common optional fields:

FieldDefaultMeaning
model_config_path""TAO experiment spec path. Required only when model_path is a TAO checkpoint.
batch_size64Number of images processed in parallel. Lower it if the GPU runs out of memory.

The input parquet must contain a filepath column. Any additional columns are carried through to the output verbatim, so metadata such as label survives into the embedding parquet.

The default template is assets/default_image_embeddings.yaml.

Encoder Consistency

When embeddings feed a mining step, every parquet compared against another must be produced with the same model and model_path. Embedding dimensionality follows the encoder — 768 for the SigLIP default, 512 for CLIP ViT-B/32 — and nothing in the output parquet records which encoder wrote it. Embeddings from different encoders are not comparable, and mismatched encoders are the most common cause of mining output that looks unrelated to the targets. Reuse one spec across every parquet in a mining run and override only input_parquet / output_parquet.

Quick Start

Run from the tao-skill-bank repo root.

Write the spec beside the output parquet. The run does not retain it, so the embeddings otherwise carry no record of the encoder that produced them. That matters here more than elsewhere: every parquet compared against another in a mining step must come from the same model and model_path, and mismatched encoders are the usual cause of mining output that looks unrelated to its targets.

bash
OUT_DIR=/absolute/path/for/this/run              # where output_parquet is written
SPEC="$OUT_DIR/image_embeddings.yaml"            # spec lives beside its output
RUN_ROOT=/absolute/path/that/contains/parquets/images/and/results
GPU_COUNT=1

python3 skills/data/tao-generate-image-embeddings/scripts/verify_image_embeddings_spec.py \
  --spec "$SPEC"

DS_IMAGE=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-data-services  # versions-key: images.tao_toolkit.data_services

docker run --rm --gpus "$GPU_COUNT" --shm-size=8g --network=host \
  -v "$RUN_ROOT:$RUN_ROOT" \
  -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
  -w "$RUN_ROOT" \
  "$DS_IMAGE" \
  embedding image_embeddings -e "$SPEC"

Do not pass --user $(id -u):$(id -g) to the TAO data-services container; the image imports transformers at startup, which calls getpass.getuser() and fails when the UID is not present in /etc/passwd.

To embed several parquets with one encoder, reuse the same spec and override the two paths per run:

bash
docker run --rm --gpus "$GPU_COUNT" --shm-size=8g --network=host \
  -v "$RUN_ROOT:$RUN_ROOT" -w "$RUN_ROOT" "$DS_IMAGE" \
  embedding image_embeddings -e "$SPEC" \
  input_parquet=/abs/path/other_input.parquet \
  output_parquet=/abs/path/other_output.parquet
Show full SKILL.md (346 more words)Show less

Generate A Spec

If the user provides parquet paths and an encoder instead of a ready spec, copy the template and fill in the nulls. Every tuning value it already carries is the one this stage wants — change one only deliberately.

bash
cp skills/data/tao-generate-image-embeddings/assets/default_image_embeddings.yaml "$SPEC"

Fill input_parquet, output_parquet and — if not using the default encoder — model and model_path, all as absolute paths, then validate:

bash
python3 skills/data/tao-generate-image-embeddings/scripts/verify_image_embeddings_spec.py --spec "$SPEC"
yaml
input_parquet: /absolute/path/filepaths.parquet
output_parquet: /absolute/path/results/embeddings.parquet
model: SigLIP
model_path: google/siglip-base-patch16-224
model_config_path: ""          # required only when model_path is a TAO .pth/.ckpt
batch_size: 64

The template is the only place a default value lives, so nothing can disagree with it. verify reports the encoder, since embeddings are only comparable to others produced by the same model and model_path.

Keep the spec, input parquet, image files, and output directory under RUN_ROOT so the same paths resolve inside the container.

Preflight

  1. Verify Docker and GPU access:
bash
docker info > /dev/null
nvidia-smi -L
  1. Resolve and pull the data-services image if needed:
bash
DS_IMAGE=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-data-services  # versions-key: images.tao_toolkit.data_services
docker image inspect "$DS_IMAGE" > /dev/null || docker pull "$DS_IMAGE"
  1. Validate the spec:
bash
python3 skills/data/tao-generate-image-embeddings/scripts/verify_image_embeddings_spec.py \
  --spec "$SPEC"
  1. Confirm RUN_ROOT contains the spec, the input parquet, the image files its filepath column points at, and the output directory. Mount RUN_ROOT to the same absolute path inside Docker.

Outputs

ArtifactLocation
embedding parquetoutput_parquet

The output parquet contains filepath, an embedding column of list-like vectors, and every extra column carried through from the input. Print its row count and column list after the run so the caller can confirm the embedding column exists.

Troubleshooting

The subtask image_embeddings requires -e/--experiment_spec_file: rerun with embedding image_embeddings -e "$SPEC".

Input parquet or images not found inside Docker: the filepath values are read verbatim. Use a RUN_ROOT mount where host and container paths are identical, and confirm the images themselves are under that mount — not just the parquet.

Model loading error with a .pth / .ckpt model_path: TAO checkpoints need model_config_path set to the training spec so the architecture can be rebuilt. HuggingFace ids and snapshot directories do not.

CUDA out of memory: lower batch_size (try 32 or 16).

Mined results look unrelated downstream: the parquets compared during mining were embedded with different encoders. Re-embed them with one shared spec — see ## Encoder Consistency.

No GPU available: embedding requires at least one CUDA GPU. Check nvidia-smi -L, the Docker --gpus flag, and the NVIDIA container toolkit installation.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references, assets) in skills/tao-generate-image-embeddings of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • assets/default_image_embeddings.yaml
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/skill_info.yaml
  • scripts/verify_image_embeddings_spec.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

TAO Image Embeddings next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

TAO Image Embeddings compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
TAO Image Embeddings this skillNVIDIA/skills3.5k—~2kAutomated safety check: NotesApache-2.0
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Optimize For GPUK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: PassMIT
Convergence TestAMD-AGI/Primus131—~2.1kAutomated safety check: PassCustom licence

Similar skills

  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Optimize For GPU

    K-Dense-AI/scientific-agent-skills

    GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Data & AnalyticsAuto-check passed
  • Convergence Test

    AMD-AGI/Primus

    Run, monitor, stop and report Primus convergence tests -- training a model on a real corpus and checking that the loss curve is healthy -- from a plain-language request such as "run convergence test…

    131 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about TAO Image Embeddings

What does TAO Image Embeddings do?

Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining. The skill takes a spec with an input parquet, an output parquet, a model of CLIP or SigLIP and a matching model path, which can be a Hugging Face model ID, a local snapshot directory or a TAO checkpoint, and runs the embedding command inside the TAO Data Services container. The input parquet needs a filepath column, and other columns such as labels are carried into the output.

When should I use TAO Image Embeddings?

TAO Image Embeddings fits situations like: embedding a folder of images before nearest-neighbor mining; generating SigLIP or CLIP embeddings for a dataset listed in a parquet file; checking that a model and model path pair match before launching a run; producing embeddings from a TAO checkpoint.

How do I install TAO Image Embeddings in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-generate-image-embeddings -a claude-code`. Or copy the skill folder (skills/tao-generate-image-embeddings in NVIDIA/skills) into .claude/skills/tao-generate-image-embeddings in your project. Claude Code loads it when a task matches its description.

How do I install TAO Image Embeddings in Codex?

Run `npx skills add NVIDIA/skills --skill tao-generate-image-embeddings -a codex`. Or copy the skill folder (skills/tao-generate-image-embeddings in NVIDIA/skills) into .agents/skills/tao-generate-image-embeddings in your project. Codex loads it when a task matches its description.

Can I use TAO Image Embeddings in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-generate-image-embeddings -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-generate-image-embeddings, .gemini/skills/tao-generate-image-embeddings, .github/skills/tao-generate-image-embeddings and .opencode/skills/tao-generate-image-embeddings in your project.

What does TAO Image Embeddings need to run?

Going by SKILL.md and its folder, TAO Image Embeddings needs Python for the scripts in its folder and the command-line tools its instructions call (docker and python3). Our summary lists: Docker with the NVIDIA container toolkit; One or more CUDA GPUs; The TAO data-services container pinned in versions.yaml. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml..

Does TAO Image Embeddings access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is TAO Image Embeddings safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does TAO Image Embeddings use?

TAO Image Embeddings is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does TAO Image Embeddings use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 138 tokens, read only when the agent opens those files.

What are the alternatives to TAO Image Embeddings?

Skills that share tags, products or a category with TAO Image Embeddings: Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 146 stars), Esmfold2 (JimLiu/science-skills, 227 stars) and Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains TAO Image Embeddings?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.