Official agent skill

Tao Mine Od Images

by NVIDIA in NVIDIA/skills

Run TAO Data Services TMM unique-neighbor matching mining from embedding parquet files for object detection workflows.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Mine Od Images

skills CLI
$ npx skills add NVIDIA/skills --skill tao-mine-od-images -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-mine-od-images --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-mine-od-images .claude/skills/tao-mine-od-images && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-mine-od-images
GitHub stars
3.6k
Token cost
~2.2k tokens
SKILL.md length
774 words
Files
9 (incl. scripts, references, assets)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run TAO Data Services TMM unique-neighbor matching mining from embedding parquet files for object detection workflows.

  • Works in 4 steps: Verify Docker and GPU access → Resolve and pull the data-services image… → Validate the spec → …
  • An object detection workflow needs to mine a bijectively-assigned set of unique source images closest to target samples
  • SKILL.md covers Inputs, Quick Start, Generate A Spec and Preflight, plus 2 more sections
  • Runs Python scripts from its folder; calls docker and python3

What it does

Tao Mine Od Images is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run TAO Data Services TMM unique-neighbor matching mining from embedding parquet files for object detection workflows. Use when an object detection workflow needs to mine a bijectively-assigned set of unique source images closest to target samples. Use global allocation when mining without class constraints. Use classstratified when rare classes are specified.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts, reference files and assets (for example `BENCHMARK.md`, `assets/default_unique_neighbor_matching.yaml` and `config/skillspector-baseline.yaml`). Compatibility notes: Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.

It sits in AI & LLM Engineering, covering DataFrames, Embeddings and Computer vision. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • An object detection workflow needs to mine a bijectively-assigned set of unique source images closest to target samples
  • Tasks that involve DataFrames
  • Tasks that involve Embeddings

Example prompts

  • “/tao-mine-od-images”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Verify Docker and GPU access
  2. Resolve and pull the data-services image if needed
  3. Validate the spec
  4. Confirm RUN_ROOT contains the spec, both input parquets (or directories), and the output directory. Mount RUN_ROOT to the same absolute…

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • docker
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Mine Od Images loads about 2.2k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 96 tokens; SKILL.md has 774 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 774 words, ~2,179 tokens.

Download SKILL.mdSave it as .claude/skills/tao-mine-od-images/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
tao-mine-od-images
description
Run TAO Data Services TMM unique-neighbor matching mining from embedding parquet files for object detection workflows. Use when an object detection workflow needs to mine a bijectively-assigned set of unique source images closest to target samples. Use global allocation when mining without class constraints. Use class_stratified when rare classes are specified.
allowed-tools
Read, Bash
compatibility
Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
tao, data, mining, object-detection, unique-neighbors, tmm

TAO Mine OD Images (Unique Neighbor Matching)

Use this skill to run TAO Data Services TMM unique-neighbor matching mining for object detection. The skill consumes pre-embedded source and target parquets and writes a directory of outputs including final_unique_files.parquet and summary.json. It does not compute embeddings; upstream steps must produce the source and target embedding parquets first.

The container entrypoint is:

bash
tmm unique_neighbor_matching -e /absolute/path/to/unique_neighbor_matching.yaml

Inputs

The user can provide either an existing spec or the fields needed to generate one.

Required spec fields:

FieldMeaning
source_pathAbsolute path to the source embeddings parquet or directory of parquets.
target_pathAbsolute path to the target embeddings parquet or directory of parquets.
output_dirAbsolute path to the output directory. Writes final_unique_files.parquet, summary.json, and per-iteration parquets.
desired_unique_countTotal number of unique source files to retrieve.

Common optional fields:

FieldDefaultMeaning
allocation_policyglobalglobal or class_stratified.
distance_metriceuclideanOne of euclidean, cosine, or manhattan. Embeddings are L2-normalized before search.
candidate_expansion_factor5Candidate-pool multiplier per iteration. Increase if desired count is not reached.
source_embedding_columnembeddingEmbedding column in source_path.
target_embedding_columnembeddingEmbedding column in target_path.
source_filepath_columnfilepathFilepath column in source_path; also the column of final_unique_files.parquet.
target_filepath_columnfilepathFilepath column in target_path.
exclude_pathnullParquet with a filepath column; those images are removed from the source pool.
source_detection_filenullCOCO .json or KITTI label directory for the source. Required for class_stratified.
target_detection_filenullCOCO .json or KITTI label directory for the target. Required for class_stratified.
detection_formatnullcoco or kitti. Required whenever a detection file is set; never inferred from the path.
rare_class_list""Comma-separated rare class names, e.g. "person,bicycle". Required for class_stratified.
save_embeddingsfalseInclude embeddings in per-iteration parquet outputs.
visualizefalseSave per-class visualization grids (requires Pillow and matplotlib).

Both input parquets must contain the filepath and embedding columns. Source and target embeddings must have been produced by the same encoder; mismatched encoders produce garbage output.

The default template is assets/default_unique_neighbor_matching.yaml.

Quick Start

Run from the tao-skill-bank repo root. Resolve the pinned TAO Data Services image from versions.yaml, verify the spec, mount the run root with identical host/container paths, and stream the Docker logs.

Write the spec into the output directory. The run does not retain it, so a mined set otherwise carries no record of the budget, allocation policy or rare-class list that produced it — and those decide which images were selected. Keeping them together makes the selection recoverable from the run alone.

bash
OUTPUT_DIR=/absolute/path/for/this/run           # output_dir in the spec
SPEC="$OUTPUT_DIR/unique_neighbor_matching.yaml" # spec lives beside its outputs
RUN_ROOT=/absolute/path/that/contains/specs/data/and/results
GPU_COUNT=1

python3 skills/data/tao-mine-od-images/scripts/verify_unique_neighbor_matching_spec.py \
  --spec "$SPEC"

DS_IMAGE=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-data-services  # versions-key: images.tao_toolkit.data_services

docker run --rm --gpus "$GPU_COUNT" --shm-size=8g --network=host \
  -v "$RUN_ROOT:$RUN_ROOT" \
  -w "$RUN_ROOT" \
  "$DS_IMAGE" \
  tmm unique_neighbor_matching -e "$SPEC"

Do not pass --user $(id -u):$(id -g) to the TAO data-services container; some TAO DS images call getpass.getuser() at startup and fail when the UID is not in /etc/passwd.

Show full SKILL.md (349 more words)Show less

Generate A Spec

If the user provides source/target paths and an output directory instead of a ready spec, copy the template and fill in the nulls. Every tuning value it already carries is the one this stage wants — change one only deliberately.

bash
cp skills/data/tao-mine-od-images/assets/default_unique_neighbor_matching.yaml "$SPEC"

Fill source_path, target_path, output_dir and desired_unique_count, all as absolute paths, then validate:

bash
python3 skills/data/tao-mine-od-images/scripts/verify_unique_neighbor_matching_spec.py --spec "$SPEC"
yaml
source_path: /absolute/path/source_embeddings.parquet
target_path: /absolute/path/target_embeddings.parquet
output_dir: /absolute/path/results/mining_output
desired_unique_count: 500
allocation_policy: global          # or class_stratified — see below
distance_metric: euclidean

For class-stratified mode set allocation_policy: class_stratified and supply rare_class_list, source_detection_file, target_detection_file and detection_format. verify rejects the policy without them: absent those fields the miner falls back to a global match, which mines the wrong images rather than failing.

The template is the only place a default value lives, so nothing can disagree with it. verify reports the budget, policy and metric, since the mined parquet is a list of filepaths and records nothing about why those files were chosen.

Keep the spec, input parquets, and output directory under RUN_ROOT so the same paths resolve inside the container.

Preflight

Before launching Docker:

  1. Verify Docker and GPU access:
bash
docker info > /dev/null
nvidia-smi -L
  1. Resolve and pull the data-services image if needed:
bash
DS_IMAGE=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-data-services  # versions-key: images.tao_toolkit.data_services
docker image inspect "$DS_IMAGE" > /dev/null || docker pull "$DS_IMAGE"
  1. Validate the spec:
bash
python3 skills/data/tao-mine-od-images/scripts/verify_unique_neighbor_matching_spec.py \
  --spec "$SPEC"
  1. Confirm RUN_ROOT contains the spec, both input parquets (or directories), and the output directory. Mount RUN_ROOT to the same absolute path inside Docker.

Outputs

ArtifactLocation
Mined source filepathsoutput_dir/final_unique_files.parquet
Coverage and allocation statsoutput_dir/summary.json
Per-iteration intermediatesoutput_dir/<subset>_iteration_<N>_topn_<K>.parquet
Per-class viz gridsoutput_dir/*.png (only if visualize: true)

final_unique_files.parquet contains one filepath column. summary.json includes retrieved_unique_count, coverage_pct, and (when detection files are provided) per-class breakdowns for the target and selected source sets.

Troubleshooting

The subtask unique_neighbor_matching requires -e/--experiment_spec_file: rerun with tmm unique_neighbor_matching -e "$SPEC".

Input path not found inside Docker: use a RUN_ROOT mount where host and container paths are identical.

ValueError: detection_format is required: set detection_format: coco or detection_format: kitti whenever source_detection_file or target_detection_file is set.

ValueError: rare_class_list is required when allocation_policy is class_stratified: set rare_class_list and both detection files when using class_stratified.

Low coverage_pct in summary.json: the source pool is smaller than desired_unique_count. Expand the pool or increase candidate_expansion_factor.

No GPU or cuDF/cuML errors: mining requires at least one CUDA GPU. Check nvidia-smi -L, the Docker --gpus flag, and the NVIDIA container toolkit installation.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references, assets) in skills/tao-mine-od-images of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • assets/default_unique_neighbor_matching.yaml
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/skill_info.yaml
  • scripts/verify_unique_neighbor_matching_spec.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Mine Od Images next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Mine Od Images compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Mine Od Images this skillNVIDIA/skills3.6k—~2.2kAutomated safety check: NotesApache-2.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Scholar Computejoshzyj/open-scholar-skill168—~15kAutomated safety check: PassCustom licence
Pi AgentK-Dense-AI/scientific-agent-skills48k1 repos~2.1kAutomated safety check: PassMIT
On Device AIsoftware-mansion-labs/skills291—~2.3kAutomated safety check: PassNone
Multimodal Dataprep Devopen-edge-platform/edge-ai-libraries171—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Scholar Compute

    joshzyj/open-scholar-skill

    Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction…

    168 GitHub stars~15k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Pi Agent

    K-Dense-AI/scientific-agent-skills

    Builds with and operates Pi, the minimal terminal coding harness.

    48k GitHub starsUsed in 1 repo~2.1k tokens
    AI & LLM EngineeringAuto-check passed
  • On Device AI

    software-mansion-labs/skills

    Build on-device AI features in React Native and Expo apps with React Native ExecuTorch.

    291 GitHub stars~2.3k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Multimodal Dataprep Dev

    open-edge-platform/edge-ai-libraries

    Develop and debug the Multimodal DataPrep microservice: its FastAPI media endpoints, in-process embedding pipeline, batch jobs, object detection, telemetry and Metrics Manager publishing, and…

    171 GitHub stars~1.3k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated 2 days ago
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated 2 days ago
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated 2 days ago
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check: notes

Questions about Tao Mine Od Images

What does Tao Mine Od Images do?

Run TAO Data Services TMM unique-neighbor matching mining from embedding parquet files for object detection workflows. Tao Mine Od Images is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run TAO Data Services TMM unique-neighbor matching mining from embedding parquet files for object detection workflows.

When should I use Tao Mine Od Images?

Tao Mine Od Images fits situations like: an object detection workflow needs to mine a bijectively-assigned set of unique source images closest to target samples; tasks that involve DataFrames; tasks that involve Embeddings.

How do I install Tao Mine Od Images in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-mine-od-images -a claude-code`. Or copy the skill folder (skills/tao-mine-od-images in NVIDIA/skills) into .claude/skills/tao-mine-od-images in your project. Claude Code loads it when a task matches its description.

How do I install Tao Mine Od Images in Codex?

Run `npx skills add NVIDIA/skills --skill tao-mine-od-images -a codex`. Or copy the skill folder (skills/tao-mine-od-images in NVIDIA/skills) into .agents/skills/tao-mine-od-images in your project. Codex loads it when a task matches its description.

Can I use Tao Mine Od Images in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-mine-od-images -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-mine-od-images, .gemini/skills/tao-mine-od-images, .github/skills/tao-mine-od-images and .opencode/skills/tao-mine-od-images in your project.

What does Tao Mine Od Images need to run?

Going by SKILL.md and its folder, Tao Mine Od Images needs Python for the scripts in its folder and the command-line tools its instructions call (docker and python3). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml..

Does Tao Mine Od Images access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tao Mine Od Images safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Tao Mine Od Images use?

Tao Mine Od Images is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Mine Od Images use?

About 2.2k tokens (SKILL.md is roughly 8.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 225 tokens, read only when the agent opens those files.

What are the alternatives to Tao Mine Od Images?

Skills that share tags, products or a category with Tao Mine Od Images: CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Scholar Compute (joshzyj/open-scholar-skill, 168 stars), Pi Agent (K-Dense-AI/scientific-agent-skills, 48k stars) and On Device AI (software-mansion-labs/skills, 291 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Mine Od Images?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.