Official agent skill

Tao Generate Referring Expressions

by NVIDIA in NVIDIA/skills

Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified…

OfficialApache-2.0Auto-check: notes

Install Tao Generate Referring Expressions

skills CLI
$ npx skills add NVIDIA/skills --skill tao-generate-referring-expressions -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-generate-referring-expressions --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-generate-referring-expressions .claude/skills/tao-generate-referring-expressions && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-generate-referring-expressions
GitHub stars
3.6k
Token cost
~2.6k tokens
SKILL.md length
990 words
Files
9 (incl. references)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified…

  • Works in 6 steps: Images: Ask for data.image_dir, the… → KITTI labels: Ask for… → Resume from existing annotations: If the… → …
  • The user wants to generate referring-expression annotations from images with KITTI labels
  • SKILL.md covers Purpose, Pipeline Architecture, Instructions and Configuration, plus 3 more sections
  • Reaches inference-api.nvidia.com; needs GOOGLE_API_KEY

What it does

Tao Generate Referring Expressions is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified expressions via VLM distillation. Use when the user wants to generate referring-expression annotations from images with KITTI labels, build region descriptions, produce grouped grounding phrases tied to bboxes, run a double-check verification pass on grounding expressions, auto-label traffic / scene images for referring…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).

The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user wants to generate referring-expression annotations from images with KITTI labels
  • Build region descriptions
  • Produce grouped grounding phrases tied to bboxes
  • Run a double-check verification pass on grounding expressions

Example prompts

  • “referring expression”
  • “region description”
  • “KITTI labels”
  • “/tao-generate-referring-expressions”

Requirements

  • Docker
  • A credential in GOOGLE_API_KEY
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).
  • Pre-approved tools (allowed-tools): Read, Bash, Write

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Images: Ask for data.image_dir, the directory containing .jpg, .jpeg, or .png images.
  2. KITTI labels: Ask for data.kitti_label_dir, the directory containing one .txt label file per image. Each label line must use KITTI format…
  3. Resume from existing annotations: If the user already has a unified annotations.jsonl from a previous run, set…
  4. API access: Ask the user which VLM endpoint they want to use. Present these five options and act on the choice
  5. Workflow steps: Choose one of
  6. Output format: Choose one of

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • inference-api.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Generate Referring Expressions loads about 2.6k tokens when it runs, and up to ~7.9k if it reads all its reference files. Until then it costs about 198 tokens; SKILL.md has 990 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~198
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 990 words, ~2,586 tokens.

Download SKILL.mdSave it as .claude/skills/tao-generate-referring-expressions/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
tao-generate-referring-expressions
description
Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified expressions via VLM distillation. Use when the user wants to generate referring-expression annotations from images with KITTI labels, build region descriptions, produce grouped grounding phrases tied to bboxes, run a double-check verification pass on grounding expressions, auto-label traffic / scene images for referring datasets, or run the image_referring_expression pipeline. Triggers include 'referring expression', 'region description', 'KITTI labels', 'spatial relationship annotation', 'auto-label image referring expression', 'image_referring_expression'.
allowed-tools
Read, Bash, Write
compatibility
Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
image, referring-expression, kitti, bounding-boxes, auto-label, vlm

Image Referring Expression Pipeline

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Generate referring-expression and grounding annotations from images with KITTI-format bounding box labels. A single VLM (Gemini or any OpenAI-compatible endpoint) runs four steps: per-object region descriptions, holistic image captions, grouped grounding expressions tied to bboxes, and an optional double-check verification pass.

Purpose

Transform (image, KITTI labels) pairs into a unified annotations.jsonl containing rich, grounded referring expressions. The VLM acts as a "teacher" annotator: Steps 0-1 see the image; Step 2 groups Step 0 outputs into grouping phrases with bbox lists; Step 3 (optional) re-examines those bboxes against the image and corrects mismatches.

Pipeline Architecture

Step 0: Region expression  ──┐
                              ├──▶  Step 2: Grounding expression  ──▶  [Step 3: Double check]
Step 1: Image caption  ──────┘                                                   (optional)
  • Step 0 (region_expr) — VLM emits one short discriminative phrase per KITTI bbox (bbox_2d, type, color, description).
  • Step 1 (image_caption) — VLM emits a holistic, location-agnostic scene caption.
  • Step 2 (grounding_expr) — VLM groups Step 0 objects into grouping phrases and returns one bbox list per group, optionally using Step 1's caption as extra context.
  • Step 3 (double_check) — VLM re-checks each Step 2 bbox against the image; bad matches are removed, slightly-off boxes get tightened.

Steps 0 and 1 run in parallel within a single thread pool (they only depend on the seed records). Each step writes its own step_<N>_*/annotations.jsonl and skips already-processed images on re-run unless workflow.force_reprocess: true.

Instructions

Initial setup

When a user wants to run this pipeline, walk through these steps:

  1. Images: Ask for data.image_dir, the directory containing .jpg, .jpeg, or .png images.

  2. KITTI labels: Ask for data.kitti_label_dir, the directory containing one .txt label file per image. Each label line must use KITTI format: <type> <truncated> <occluded> <alpha> <bbox_left> <bbox_top> <bbox_right> <bbox_bottom> .... Lines with fewer than 8 fields are silently skipped. Set this even for Step 1-only runs because Steps 0 and 2 require it.

  3. Resume from existing annotations: If the user already has a unified annotations.jsonl from a previous run, set data.input_annotations_jsonl to that file instead of seeding from data.image_dir and data.kitti_label_dir.

  4. API access: Ask the user which VLM endpoint they want to use. Present these five options and act on the choice:

    1. Gemini — set vlm.backend: "gemini"; require GOOGLE_API_KEY (env var or vlm.gemini.api_key).
    2. NIM (e.g. https://inference-api.nvidia.com/v1) — set vlm.backend: "openai"; collect base_url, model_name, and api_key.
    3. TAO inference microservice (self-hosted, OpenAI-compatible). Confirm whether the server is already running:
      • Running — collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
      • Not running — guide the user through the skills/applications/tao-run-inference-service skill, which stands up a local TAO inference microservice with an OpenAI-compatible API. Before promising a specific model, check skills/applications/tao-run-inference-service/references/service.yaml for valid_network_arch_config_basenames. Once the server is up, collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
    4. vLLM (self-hosted, OpenAI-compatible). Confirm whether the server is already running:
      • Running — collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
      • Not running — follow references/vllm_server.md to install and launch a vLLM server, then collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
    5. Custom (any other OpenAI-compatible endpoint) — set vlm.backend: "openai"; collect base_url, model_name, and (optionally) api_key.

    If the user has no endpoint and does not want to set one up, stop and help resolve API access first.

  5. Workflow steps: Choose one of:

    • Full pipeline: ["0", "1", "2", "3"]
    • No caption generation: ["0", "2", "3"], where Step 2 falls back to image-only context
    • No verification: ["0", "1", "2"]
    • Custom subset: any supported subset of steps
  6. Output format: Choose one of:

    • jsonl: unified schema only
    • legacy: byte-compatible .txt.stepN files only
    • both: writes both formats and is the default for downstream tooling
Show full SKILL.md (400 more words)Show less
Running the pipeline

The pipeline runs inside the TAO Toolkit container via the auto_label CLI:

bash
auto_label generate -e /path/to/spec.yaml \
    results_dir=/results \
    image_referring_expression.data.image_dir=/data/images \
    image_referring_expression.data.kitti_label_dir=/data/labels \
    image_referring_expression.vlm.gemini.api_key=$GOOGLE_API_KEY

Generate a default spec: auto_label default_specs results_dir=/results module_name=auto_label, then set autolabel_type: "image_referring_expression". All fields support Hydra dot-notation overrides on the command line.

See references/configuration.md for the full YAML structure, all parameters, model/endpoint setup, and error patterns.

  1. Run on 5-10 images with all four steps.
  2. Inspect step_0_region_expr/annotations.jsonl — are object types, colors, and discriminating phrases accurate?
  3. Inspect step_2_grounding_expr/annotations.jsonl — are objects grouped sensibly, and do bbox coordinates match the described groups?
  4. Inspect step_3_double_check/annotations.jsonl — were mismatched bboxes removed or tightened? Are any new errors introduced (rare)?
  5. If quality is insufficient, switch the VLM to a stronger model (e.g. gemini-2.5-pro or a larger Qwen3-VL endpoint), raise media_resolution / max_output_tokens, then re-run with workflow.force_reprocess=true.
  6. Scale to the full dataset once satisfied.

Configuration

Key configuration fields (full reference in references/configuration.md):

FieldDefaultDescription
workflow.steps["0","1","2","3"]Which steps to execute (0=region_expr, 1=image_caption, 2=grounding_expr, 3=double_check)
workflow.max_workers4Parallel threads per step (watch API rate limits)
workflow.force_reprocessfalseIgnore cached per-step outputs and reprocess from scratch
workflow.output_format"jsonl" (set to "both" in the default spec)"jsonl", "legacy", or "both"
vlm.backend"gemini""gemini" or "openai" (OpenAI-compatible endpoint)
data.image_dirrequiredDirectory of input images (.jpg / .jpeg / .png)
data.kitti_label_dirrequired (unless resuming)Directory of KITTI-format .txt label files
data.input_annotations_jsonl""Optional pre-seeded annotations.jsonl (skips KITTI seeding)

Inputs

Two ways to seed the pipeline:

  1. Image directory + KITTI labels (default). Set data.image_dir and data.kitti_label_dir. The orchestrator walks the image directory, reads the matching <stem>.txt KITTI file, parses bboxes (fields 0 + 4-7), reads each image's width/height via PIL, and writes a seed_annotations.jsonl to results_dir/.
  2. Pre-seeded annotations JSONL (resume / pre-computed regions). Set data.input_annotations_jsonl to a file with one {"image_id", "image_path", "width", "height", "kitti_bboxes": [...]} object per line.

Outputs

All outputs go to results_dir/:

  • seed_annotations.jsonl — initial per-image records (unless input_annotations_jsonl was supplied).
  • step_0_region_expr/annotations.jsonl — adds regions[] (each with bbox/bbox_2d, type, color, description).
  • step_1_image_caption/annotations.jsonl — adds caption (string).
  • step_2_grounding_expr/annotations.jsonl — adds expressions[] (each {text, instances: [{bbox: [x1,y1,x2,y2]}]}).
  • step_3_double_check/annotations.jsonl — same shape as Step 2, with bboxes removed/updated.
  • results_dir/annotations.jsonl — copy of the last completed step's output.
  • When workflow.output_format is "legacy" or "both", each step also writes byte-compatible step_<N>_*/labels/<stem>.txt.stepN files for the original 2d-data-engine tooling.

Prerequisites

  • Container: nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt <!-- versions-key: images.tao_toolkit.pyt -->
  • API access: At least one VLM endpoint (Gemini API key or OpenAI-compatible endpoint capable of image input)
  • PIL / Pillow: Required to read image dimensions during seeding (already present in the TAO container)

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/tao-generate-referring-expressions of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/configuration.md
  • references/skill_info.yaml
  • references/vllm_server.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Generate Referring Expressions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Generate Referring Expressions compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Generate Referring Expressions this skillNVIDIA/skills3.6k—~2.6kAutomated safety check: NotesApache-2.0
Reference Design Contractnexu-io/open-design100k—~1.3kAutomated safety check: PassApache-2.0
Makepad Referencesickn33/agentic-awesome-skills47k2 repos~573Automated safety check: PassMIT
Matematico Taosickn33/agentic-awesome-skills47k2 repos~459Automated safety check: PassMIT
Turn This Into A Threadsickn33/agentic-awesome-skills47k1 repos~1.4kAutomated safety check: PassMIT
Write API Referencevercel/next.js143k—~2.2kAutomated safety check: PassMIT

Similar skills

  • Reference Design Contract

    nexu-io/open-design

    Turn vague taste, screenshots, URLs, product notes, or "make it feel like this" references into a grounded DESIGN.md plus an implementation handoff.

    100k GitHub stars~1.3k tokensUpdated yesterday
    Frontend & DesignAuto-check passed
  • Makepad Reference

    sickn33/agentic-awesome-skills

    This category provides reference materials for debugging, code quality, and advanced layout patterns.

    47k GitHub starsUsed in 2 repos~573 tokens
    DevelopmentAuto-check passed
  • Matematico Tao

    sickn33/agentic-awesome-skills

    Matemático ultra-avançado inspirado em Terence Tao. An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~459 tokens
    Auto-check passed
  • Turn This Into A Thread

    sickn33/agentic-awesome-skills

    Turn a long idea, transcript, article or draft into a Twitter/X thread where every tweet stands alone.

    47k GitHub starsUsed in 1 repo~1.4k tokens
    Writing & ContentAuto-check passed
  • Write API Reference

    vercel/next.js

    Official

    Produces API reference documentation for Next.js APIs: functions, components, file conventions, directives, and config options.

    143k GitHub stars~2.2k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Reference List Builder

    davila7/claude-code-templates

    Format professional references properly and prepare reference materials.

    33k GitHub starsUsed in 2 repos~2.1k tokens
    Research & ScienceAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Generate Referring Expressions

What does Tao Generate Referring Expressions do?

Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified…. Tao Generate Referring Expressions is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Four-step image referring-expression pipeline: turns images plus KITTI bounding-box labels into region descriptions, scene captions, grounded referring expressions, and (optionally) verified expressions via VLM distillation.

When should I use Tao Generate Referring Expressions?

Tao Generate Referring Expressions fits situations like: the user wants to generate referring-expression annotations from images with KITTI labels; build region descriptions; produce grouped grounding phrases tied to bboxes; run a double-check verification pass on grounding expressions.

How do I install Tao Generate Referring Expressions in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-generate-referring-expressions -a claude-code`. Or copy the skill folder (skills/tao-generate-referring-expressions in NVIDIA/skills) into .claude/skills/tao-generate-referring-expressions in your project. Claude Code loads it when a task matches its description.

How do I install Tao Generate Referring Expressions in Codex?

Run `npx skills add NVIDIA/skills --skill tao-generate-referring-expressions -a codex`. Or copy the skill folder (skills/tao-generate-referring-expressions in NVIDIA/skills) into .agents/skills/tao-generate-referring-expressions in your project. Codex loads it when a task matches its description.

Can I use Tao Generate Referring Expressions in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-generate-referring-expressions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-generate-referring-expressions, .gemini/skills/tao-generate-referring-expressions, .github/skills/tao-generate-referring-expressions and .opencode/skills/tao-generate-referring-expressions in your project.

What does Tao Generate Referring Expressions need to run?

Going by SKILL.md and its folder, Tao Generate Referring Expressions needs credentials named GOOGLE_API_KEY. Our summary lists: Docker; A credential in GOOGLE_API_KEY. Its frontmatter pre-approves these tools: Read, Bash, Write. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible)..

Does Tao Generate Referring Expressions access the network?

SKILL.md names 1 domain. In commands or code: inference-api.nvidia.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Tao Generate Referring Expressions safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Generate Referring Expressions use?

Tao Generate Referring Expressions is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Generate Referring Expressions use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Tao Generate Referring Expressions?

Skills that share tags, products or a category with Tao Generate Referring Expressions: Reference Design Contract (nexu-io/open-design, 100k stars), Makepad Reference (sickn33/agentic-awesome-skills, 47k stars), Matematico Tao (sickn33/agentic-awesome-skills, 47k stars) and Turn This Into A Thread (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Generate Referring Expressions?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.