Official agent skill

Tao Mine Nearest Neighbors

by NVIDIA in NVIDIA/skills

Run TAO Data Services TMM nearest-neighbor mining from embedding parquet files.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Mine Nearest Neighbors

skills CLI
$ npx skills add NVIDIA/skills --skill tao-mine-nearest-neighbors -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-mine-nearest-neighbors --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-mine-nearest-neighbors .claude/skills/tao-mine-nearest-neighbors && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-mine-nearest-neighbors
GitHub stars
3.6k
Token cost
~1.6k tokens
SKILL.md length
590 words
Files
10 (incl. scripts, references, assets)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run TAO Data Services TMM nearest-neighbor mining from embedding parquet files.

  • Works in 4 steps: Verify Docker and GPU access → Resolve and pull the data-services image… → Validate the spec → …
  • A workflow needs to mine source samples closest to target samples
  • SKILL.md covers Inputs, Quick Start, Generate A Spec and Preflight, plus 2 more sections
  • Runs Python scripts from its folder; calls docker and python3

What it does

Tao Mine Nearest Neighbors is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run TAO Data Services TMM nearest-neighbor mining from embedding parquet files. Use when a workflow needs to mine source samples closest to target samples.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts, reference files and assets (for example `BENCHMARK.md`, `assets/default_nearest_neighbors.yaml` and `config/skillspector-baseline.yaml`). Compatibility notes: Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.

It sits in AI & LLM Engineering, covering DataFrames and Embeddings. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • A workflow needs to mine source samples closest to target samples
  • Tasks that involve DataFrames
  • Tasks that involve Embeddings

Example prompts

  • “/tao-mine-nearest-neighbors”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.
  • Pre-approved tools (allowed-tools): Read, Bash

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Verify Docker and GPU access
  2. Resolve and pull the data-services image if needed
  3. Validate the spec
  4. Confirm RUN_ROOT contains the spec, both input parquets, and the output directory. Mount RUN_ROOT to the same absolute path inside Docker.

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • docker
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Mine Nearest Neighbors loads about 1.6k tokens when it runs, and up to ~1.7k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 590 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 590 words, ~1,565 tokens.

Download SKILL.mdSave it as .claude/skills/tao-mine-nearest-neighbors/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
tao-mine-nearest-neighbors
description
Run TAO Data Services TMM nearest-neighbor mining from embedding parquet files. Use when a workflow needs to mine source samples closest to target samples.
allowed-tools
Read, Bash
compatibility
Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
tao, data, mining, nearest-neighbors, tmm

TAO Mine Nearest Neighbors

Use this skill to run TAO Data Services TMM nearest-neighbor mining. The skill consumes embedding parquets and writes a mined source-sample parquet plus a mining summary. It does not compute embeddings; upstream steps must produce the source and target embedding parquets first.

The container entrypoint is:

bash
tmm nearest_neighbors -e /absolute/path/to/nearest_neighbors.yaml

TAO Data Services requires -e/--experiment_spec_file. The tmm console script converts that YAML into Hydra --config-path and --config-name arguments internally.

Inputs

The user can provide either an existing nearest-neighbors YAML spec or the fields needed to generate one.

Required spec fields:

FieldMeaning
source_parquetAbsolute path to the candidate/source embeddings parquet.
target_parquetAbsolute path to the target/query embeddings parquet.
output_parquetAbsolute path where TAO Data Services should write mined source filepaths.

Common optional fields:

FieldDefaultMeaning
topn5Number of nearest source samples to retrieve per target sample.
knn_metriccosineOne of cosine, euclidean, or manhattan.
source_embed_column_nameembeddingEmbedding column in source_parquet.
target_embed_column_nameembeddingEmbedding column in target_parquet.
filter_by_label"false"String flag. When "true", TAO DS filters neighbors by matching label columns when both parquets provide labels.
distance_threshold-1.0Maximum distance to keep. Negative disables thresholding.

Both input parquets must contain a filepath column and a list-like embedding column. If filter_by_label is "true", both parquets should also contain label.

The default template is assets/default_nearest_neighbors.yaml.

Quick Start

Run from the tao-skill-bank repo root. Resolve the pinned TAO Data Services image from versions.yaml, verify the spec, mount the run root with identical host/container paths, and stream the Docker logs.

bash
SPEC=/absolute/path/to/nearest_neighbors.yaml
RUN_ROOT=/absolute/path/that/contains/specs/data/and/results
GPU_COUNT=1

python3 skills/data/tao-mine-nearest-neighbors/scripts/verify_nearest_neighbors_spec.py \
  --spec "$SPEC"

DS_IMAGE="$(scripts/resolve_versions_key.py images.tao_toolkit.data_services)"

docker run --rm --gpus "$GPU_COUNT" --shm-size=8g --network=host \
  -v "$RUN_ROOT:$RUN_ROOT" \
  -w "$RUN_ROOT" \
  "$DS_IMAGE" \
  tmm nearest_neighbors -e "$SPEC"

Use at least one GPU. Choose GPU_COUNT from the hardware available to the host or platform that will run the container. If the user does not know the right value, inspect the host with nvidia-smi -L or ask which GPU allocation the run should use.

Do not pass --user $(id -u):$(id -g) to the TAO data-services container unless you have verified the image supports that UID. Some TAO DS images import Python packages that call getpass.getuser() at startup and fail when the UID is not present in /etc/passwd.

Show full SKILL.md (255 more words)Show less

Generate A Spec

If the user provides source/target/output parquet paths instead of a ready spec, generate a spec from the default template:

bash
python3 skills/data/tao-mine-nearest-neighbors/scripts/prepare_nearest_neighbors_spec.py \
  --source-parquet /absolute/path/source_embeddings.parquet \
  --target-parquet /absolute/path/target_embeddings.parquet \
  --output-parquet /absolute/path/results/mined.parquet \
  --output-spec /absolute/path/specs/nearest_neighbors.yaml \
  --topn 5 \
  --knn-metric cosine \
  --filter-by-label false \
  --distance-threshold -1.0

The generated YAML uses absolute paths. Keep the spec, input parquets, and output directory under RUN_ROOT so the same paths resolve inside the container.

Preflight

Before launching Docker:

  1. Verify Docker and GPU access:
bash
docker info > /dev/null
nvidia-smi -L
  1. Resolve and pull the data-services image if needed:
bash
DS_IMAGE="$(scripts/resolve_versions_key.py images.tao_toolkit.data_services)"
docker image inspect "$DS_IMAGE" > /dev/null || docker pull "$DS_IMAGE"
  1. Validate the spec:
bash
python3 skills/data/tao-mine-nearest-neighbors/scripts/verify_nearest_neighbors_spec.py \
  --spec "$SPEC"
  1. Confirm RUN_ROOT contains the spec, both input parquets, and the output directory. Mount RUN_ROOT to the same absolute path inside Docker.

Outputs

The skill promises the artifacts named by the spec:

ArtifactLocation
mined parquetoutput_parquet
mining summarymining_summary.txt next to output_parquet

The current TAO Data Services nearest_neighbors task writes a mined parquet with unique source filepath rows. The summary file reports mining counts such as queries processed, neighbors considered, duplicates removed, and any label/distance filtering.

Troubleshooting

The subtask nearest_neighbors requires -e/--experiment_spec_file: rerun with tmm nearest_neighbors -e "$SPEC". Hydra overrides alone are not enough.

Input parquet not found inside Docker: the YAML path must be visible inside the container. Use a RUN_ROOT mount where the host and container paths are identical.

Output directory is not writable after Docker exits: the TAO DS container may have written files as root. Inform the user, report which artifacts were produced, and ask whether to repair permissions on the output directory before continuing.

No GPU or cuDF/cuML errors: nearest-neighbor mining requires at least one CUDA GPU. Check nvidia-smi -L, the Docker --gpus flag, and the NVIDIA container toolkit installation.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files (scripts, references, assets) in skills/tao-mine-nearest-neighbors of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • assets/default_nearest_neighbors.yaml
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/skill_info.yaml
  • scripts/prepare_nearest_neighbors_spec.py
  • scripts/verify_nearest_neighbors_spec.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Mine Nearest Neighbors next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Mine Nearest Neighbors compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Mine Nearest Neighbors this skillNVIDIA/skills3.6k—~1.6kAutomated safety check: NotesApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Codebase Managementgiancarloerra/SocratiCode3.3k1 repos~1.8kAutomated safety check: PassAGPL-3.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Sentence-Transformers Training Routerhuggingface/skills11k1 repos~2.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Codebase Management

    giancarloerra/SocratiCode

    Set up, index, and manage SocratiCode codebase indexing. An agent skill from giancarloerra/SocratiCode.

    3.3k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Routes a sentence-transformers training task to the right model type and required reference docs and example scripts, covering bi-encoders, rerankers, sparse and multi-vector models.

    11k GitHub starsUsed in 1 repo~2.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Mashup Mods

    rehan-remade/universal-modder

    Build cross-game mashups and total conversions, the "Minecraft inside Elden Ring" or "skateboarding in MW2" kind.

    6.5k GitHub stars~3.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Mine Nearest Neighbors

What does Tao Mine Nearest Neighbors do?

Run TAO Data Services TMM nearest-neighbor mining from embedding parquet files. Tao Mine Nearest Neighbors is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run TAO Data Services TMM nearest-neighbor mining from embedding parquet files.

When should I use Tao Mine Nearest Neighbors?

Tao Mine Nearest Neighbors fits situations like: A workflow needs to mine source samples closest to target samples; tasks that involve DataFrames; tasks that involve Embeddings.

How do I install Tao Mine Nearest Neighbors in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-mine-nearest-neighbors -a claude-code`. Or copy the skill folder (skills/tao-mine-nearest-neighbors in NVIDIA/skills) into .claude/skills/tao-mine-nearest-neighbors in your project. Claude Code loads it when a task matches its description.

How do I install Tao Mine Nearest Neighbors in Codex?

Run `npx skills add NVIDIA/skills --skill tao-mine-nearest-neighbors -a codex`. Or copy the skill folder (skills/tao-mine-nearest-neighbors in NVIDIA/skills) into .agents/skills/tao-mine-nearest-neighbors in your project. Codex loads it when a task matches its description.

Can I use Tao Mine Nearest Neighbors in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-mine-nearest-neighbors -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-mine-nearest-neighbors, .gemini/skills/tao-mine-nearest-neighbors, .github/skills/tao-mine-nearest-neighbors and .opencode/skills/tao-mine-nearest-neighbors in your project.

What does Tao Mine Nearest Neighbors need to run?

Going by SKILL.md and its folder, Tao Mine Nearest Neighbors needs Python for the scripts in its folder and the command-line tools its instructions call (docker and python3). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, one or more CUDA GPUs, and the TAO data-services container pinned in versions.yaml..

Does Tao Mine Nearest Neighbors access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tao Mine Nearest Neighbors safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Tao Mine Nearest Neighbors use?

Tao Mine Nearest Neighbors is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Mine Nearest Neighbors use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 141 tokens, read only when the agent opens those files.

What are the alternatives to Tao Mine Nearest Neighbors?

Skills that share tags, products or a category with Tao Mine Nearest Neighbors: Chroma Vector Database (Orchestra-Research/AI-Research-SKILLs, 13k stars), SageMaker Serving Image Selection (huggingface/skills, 11k stars), Codebase Management (giancarloerra/SocratiCode, 3.3k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Mine Nearest Neighbors?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.