Official agent skill

Tao Generate Image Grounding

by NVIDIA in NVIDIA/skills

Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM.

OfficialApache-2.0Auto-check: notesMedia & Creative

Install Tao Generate Image Grounding

skills CLI
$ npx skills add NVIDIA/skills --skill tao-generate-image-grounding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-generate-image-grounding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-generate-image-grounding .claude/skills/tao-generate-image-grounding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-generate-image-grounding
GitHub stars
3.5k
Token cost
~2k tokens
SKILL.md length
744 words
Files
9 (incl. references)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM.

  • Works in 5 steps: Input JSONL: Ask for the JSONL path.… → Image root: If any image_path values are… → API access: Ask the user which VLM… → …
  • The user wants to ground captions to bboxes
  • SKILL.md covers Purpose, Pipeline Architecture, Instructions and Configuration, plus 3 more sections
  • Reaches inference-api.nvidia.com; needs GOOGLE_API_KEY

What it does

Tao Generate Image Grounding is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, or run the imagegrounding pipeline. Triggers include 'image grounding', 'phrase grounding', 'ground captions', 'auto-label image grounding', 'imagegrounding'.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).

It sits in Media & Creative, covering Image generation. It works with OpenAI. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user wants to ground captions to bboxes
  • Generate phrase-grounded annotations
  • Auto-label images for grounding
  • Run the imagegrounding pipeline

Example prompts

  • “image grounding”
  • “phrase grounding”
  • “ground captions”
  • “/tao-generate-image-grounding”

Requirements

  • Docker
  • A credential in GOOGLE_API_KEY
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).
  • Pre-approved tools (allowed-tools): Read, Bash, Write

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Input JSONL: Ask for the JSONL path. Each line must be one object like {"image_path": "...", "caption": "..."}. image_path can be absolute…
  2. Image root: If any image_path values are relative, set data.image_root to the directory they should resolve from.
  3. API access: Ask the user which VLM endpoint they want to use. Present these five options and act on the choice
  4. Workflow steps: Choose one of
  5. Resume vs fresh run: By default, the workflow reuses checkpoints and skips completed records. To reprocess everything, set…

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • inference-api.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Generate Image Grounding loads about 2k tokens when it runs, and up to ~6.4k if it reads all its reference files. Until then it costs about 117 tokens; SKILL.md has 744 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 744 words, ~2,006 tokens.

Download SKILL.mdSave it as .claude/skills/tao-generate-image-grounding/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
tao-generate-image-grounding
description
Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Use when the user wants to ground captions to bboxes, generate phrase-grounded annotations, auto-label images for grounding, or run the image_grounding pipeline. Triggers include 'image grounding', 'phrase grounding', 'ground captions', 'auto-label image grounding', 'image_grounding'.
allowed-tools
Read, Bash, Write
compatibility
Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
image, grounding, bounding-boxes, auto-label, vlm, 2d-grounding

Image Grounding Pipeline

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Turn (image, caption) pairs into per-image grounded annotations: cleaned captions, referring expressions with character spans, and pixel-space bounding boxes for each expression. A single VLM (Gemini or any OpenAI-compatible endpoint) handles both steps.

Purpose

Generate phrase-grounded training data for referring-expression and grounding models. The VLM acts as a "teacher" annotator: Step 0 extracts referring expressions from the caption while looking at the image; Step 1 returns one bbox set per expression for each image.

Pipeline Architecture

Step 0: Expression extraction  → VLM cleans caption, extracts referring expressions + char spans
Step 1: Phrase grounding       → VLM returns pixel bboxes + scores per expression

Steps are individually selectable via workflow.steps. Each step writes a per-sample checkpoint to step_<N>_*/.ckpt/<sample_id>.json and skips already-processed records on re-run. Set workflow.force_reprocess: true to ignore checkpoints and reprocess from scratch.

Instructions

Initial setup

When a user wants to run this pipeline, walk through these steps:

  1. Input JSONL: Ask for the JSONL path. Each line must be one object like {"image_path": "...", "caption": "..."}. image_path can be absolute or relative.

  2. Image root: If any image_path values are relative, set data.image_root to the directory they should resolve from.

  3. API access: Ask the user which VLM endpoint they want to use. Present these five options and act on the choice:

    1. Gemini — set vlm.backend: "gemini"; require GOOGLE_API_KEY (env var or vlm.gemini.api_key).
    2. NIM (e.g. https://inference-api.nvidia.com/v1) — set vlm.backend: "openai"; collect base_url, model_name, and api_key.
    3. TAO inference microservice (self-hosted, OpenAI-compatible). Confirm whether the server is already running:
      • Running — collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
      • Not running — guide the user through the skills/applications/tao-run-inference-service skill, which stands up a local TAO inference microservice with an OpenAI-compatible API. Before promising a specific model, check skills/applications/tao-run-inference-service/references/service.yaml for valid_network_arch_config_basenames. Once the server is up, collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
    4. vLLM (self-hosted, OpenAI-compatible). Confirm whether the server is already running:
      • Running — collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
      • Not running — follow references/vllm_server.md to install and launch a vLLM server, then collect base_url, model_name, and (optionally) api_key; set vlm.backend: "openai".
    5. Custom (any other OpenAI-compatible endpoint) — set vlm.backend: "openai"; collect base_url, model_name, and (optionally) api_key.

    If the user has no endpoint and does not want to set one up, stop and help resolve API access first.

  4. Workflow steps: Choose one of:

    • Full pipeline: ["0", "1"]
    • Expression extraction only: ["0"]
    • Grounding only: ["1"], which requires existing step-0 output at results_dir/step_0_expression_extraction/annotations.jsonl
  5. Resume vs fresh run: By default, the workflow reuses checkpoints and skips completed records. To reprocess everything, set image_grounding.workflow.force_reprocess=true.

Show full SKILL.md (323 more words)Show less
Running the pipeline

The pipeline runs inside the TAO Toolkit container via the auto_label CLI:

bash
auto_label generate -e /path/to/spec.yaml \
    results_dir=/results \
    image_grounding.data.input_jsonl=/data/captions.jsonl \
    image_grounding.data.image_root=/data/images \
    image_grounding.vlm.gemini.api_key=$GOOGLE_API_KEY

Generate a default spec: auto_label default_specs results_dir=/results module_name=auto_label, then set autolabel_type: "image_grounding". All fields support Hydra dot-notation overrides on the command line.

See references/configuration.md for the full YAML structure, all parameters, model/endpoint setup, and error patterns.

  1. Run on 5-10 images with both steps
  2. Inspect step_0_expression_extraction/annotations.jsonl — are cleaned_caption and expressions[] accurate? Are the right noun phrases captured?
  3. Inspect step_1_grounding/annotations.jsonl — do the bboxes in expressions[].instances[] look right? Are confidence scores reasonable?
  4. If quality is insufficient, switch the VLM to a stronger model (e.g. gemini-2.5-pro) or raise media_resolution/max_output_tokens, then re-run with force_reprocess=true.
  5. Scale to the full dataset once satisfied.

Configuration

Key configuration fields (full reference in references/configuration.md):

FieldDefaultDescription
workflow.steps["0","1"]Which pipeline steps to execute ("0" = expressions, "1" = grounding)
workflow.max_workers4Parallel threads per step (watch API rate limits)
workflow.force_reprocessfalseIgnore per-sample checkpoints and reprocess from scratch
vlm.backend"gemini""gemini" or "openai" (OpenAI-compatible endpoint)
data.input_jsonlrequiredPath to input JSONL with image_path + caption per line
data.image_root""Optional prefix for resolving relative image_path entries

Inputs

A single JSONL file at data.input_jsonl. One JSON object per line:

FieldRequiredDescription
image_pathyesAbsolute path, or relative path resolved against data.image_root
captionyesFree-text caption for the image
image_idnoStable identifier; auto-derived from the filename if missing
width, heightnoImage dimensions in pixels; default to 1920×1080 for bbox clamping if missing

Outputs

All outputs go to results_dir/:

  • step_0_expression_extraction/annotations.jsonl — per-record output enriched with cleaned_caption and expressions[] (each with text, expression_id, char_span, noun_chunk, empty instances[]).
  • step_1_grounding/annotations.jsonl — same records with expressions[].instances[] filled in (each instance has bbox: [x1,y1,x2,y2] in pixel space, score in [0.0, 1.0], and bbox_id).
  • results_dir/annotations.jsonl — copy of the last step's output for convenience.
  • step_<N>_*/.ckpt/<sample_id>.json — per-sample checkpoints used for resume.

Prerequisites

  • Container: nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt <!-- versions-key: images.tao_toolkit.pyt -->
  • API access: At least one VLM endpoint (Gemini API key or OpenAI-compatible endpoint capable of image input)

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/tao-generate-image-grounding of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/configuration.md
  • references/skill_info.yaml
  • references/vllm_server.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Tao Generate Image Grounding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Generate Image Grounding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Generate Image Grounding this skillNVIDIA/skills3.5k—~2kAutomated safety check: NotesApache-2.0
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
GPT Image Generation CLIwuyoscar/GPT-Image2-Skill5.7k—~2.5kAutomated safety check: NotesMIT
Imagegentheowenyoung/home1154 repos~4.8kAutomated safety check: PassApache-2.0
Openai Image Gentrpc-group/trpc-agent-go1.9k12 repos~843Automated safety check: PassApache-2.0
Image Generationonyx-dot-app/onyx32k1 repos~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • GPT Image Generation CLI

    wuyoscar/GPT-Image2-Skill

    Generates and edits images with GPT Image 2 or 2.5 through a packaged CLI and a prompt gallery, after settling which model fits the request.

    5.7k GitHub stars~2.5k tokensUpdated 9 days ago
    Media & CreativeAuto-check: notes
  • Imagegen

    theowenyoung/home

    Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts.

    115 GitHub starsUsed in 4 repos~4.8k tokens
    Media & CreativeAuto-check passed
  • Openai Image Gen

    trpc-group/trpc-agent-go

    Batch-generate images via OpenAI Images API. An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 12 repos~843 tokens
    Media & CreativeAuto-check passed
  • Image Generation

    onyx-dot-app/onyx

    Generate or edit raster images (photos, illustrations, textures, sprites, mockups, logos, infographics) using the workspace's configured image-generation provider via onyx-cli image.

    32k GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • BlockRun Image Generation

    BlockRunAI/ClawRouter

    Generates or edits images through ClawRouter's local image API, with a choice of models and sizes and payment handled automatically through x402.

    6.6k GitHub stars~2.1k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Works with

Questions about Tao Generate Image Grounding

What does Tao Generate Image Grounding do?

Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM. Tao Generate Image Grounding is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Two-step image grounding pipeline: extracts referring expressions from (image, caption) pairs and grounds them to pixel-space bounding boxes via a VLM.

When should I use Tao Generate Image Grounding?

Tao Generate Image Grounding fits situations like: the user wants to ground captions to bboxes; generate phrase-grounded annotations; auto-label images for grounding; run the imagegrounding pipeline.

How do I install Tao Generate Image Grounding in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-generate-image-grounding -a claude-code`. Or copy the skill folder (skills/tao-generate-image-grounding in NVIDIA/skills) into .claude/skills/tao-generate-image-grounding in your project. Claude Code loads it when a task matches its description.

How do I install Tao Generate Image Grounding in Codex?

Run `npx skills add NVIDIA/skills --skill tao-generate-image-grounding -a codex`. Or copy the skill folder (skills/tao-generate-image-grounding in NVIDIA/skills) into .agents/skills/tao-generate-image-grounding in your project. Codex loads it when a task matches its description.

Can I use Tao Generate Image Grounding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-generate-image-grounding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-generate-image-grounding, .gemini/skills/tao-generate-image-grounding, .github/skills/tao-generate-image-grounding and .opencode/skills/tao-generate-image-grounding in your project.

What does Tao Generate Image Grounding need to run?

Going by SKILL.md and its folder, Tao Generate Image Grounding needs credentials named GOOGLE_API_KEY. Our summary lists: Docker; A credential in GOOGLE_API_KEY. Its frontmatter pre-approves these tools: Read, Bash, Write. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible)..

Does Tao Generate Image Grounding access the network?

SKILL.md names 1 domain. In commands or code: inference-api.nvidia.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Tao Generate Image Grounding safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Generate Image Grounding use?

Tao Generate Image Grounding is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Generate Image Grounding use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.4k tokens, read only when the agent opens those files.

What are the alternatives to Tao Generate Image Grounding?

Skills that share tags, products or a category with Tao Generate Image Grounding: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), GPT Image Generation CLI (wuyoscar/GPT-Image2-Skill, 5.7k stars), Imagegen (theowenyoung/home, 115 stars) and Openai Image Gen (trpc-group/trpc-agent-go, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Generate Image Grounding?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.