Official agent skill

VLM BCQ Gap Analysis

by NVIDIA in NVIDIA/skills

Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install VLM BCQ Gap Analysis

skills CLI
$ npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-analyze-gaps-vlm-bcq --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-analyze-gaps-vlm-bcq .claude/skills/tao-analyze-gaps-vlm-bcq && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-analyze-gaps-vlm-bcq
GitHub stars
3.6k
Token cost
~1.3k tokens
SKILL.md length
523 words
Files
9 (incl. scripts, references, assets)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.

  • Pulling failure cases out of a VLM yes/no evaluation
  • SKILL.md covers Purpose, Usage, Inputs and Outputs, plus 2 more sections
  • Runs Python scripts from its folder; calls python3
  • Preparing inputs for root-cause analysis in a DEFT iteration

What it does

This skill reads a predictions JSON from a vision-language model run on a binary yes/no question task, checks each response against ground truth and extracts the false positive and false negative samples. Results go to kpi_gaps.jsonl with counts in kpi_gaps_report.txt, which downstream root-cause stages (such as cosmos generation and root cause analysis) use to drive a DEFT iteration.

A bundled helper, prepare_vlm_bcq_spec.py, builds the vlm_bcq_spec.yaml with the predictions file, an optional videos directory for relative video IDs and a results folder. The vlm_bcq action then runs inside the TAO Toolkit data services container with the spec passed through -e. The platform request must ask for exactly one GPU, because the image calls nvidia-smi even though the analysis itself does no GPU compute. If the session was not set up by the TAO skill bank plugin, tao-setup runs first.

When your agent uses it

  • Pulling failure cases out of a VLM yes/no evaluation
  • Preparing inputs for root-cause analysis in a DEFT iteration
  • Getting false positive and false negative counts from a predictions JSON

Example prompts

  • “Extract the false positives and false negatives from my VLM predictions JSON.”
  • “Generate the vlm_bcq spec for results.json and run the gap analysis.”
  • “Analyze VLM BCQ gaps for these predictions, where the video IDs are relative to my videos folder.”

Requirements

  • Docker, NVIDIA Container Toolkit and one visible GPU
  • A predictions JSON from a VLM binary-classification run
  • The TAO Toolkit data services container
  • Compatibility (from SKILL.md): Requires Docker, NVIDIA Container Toolkit, and one visible GPU.
  • Pre-approved tools (allowed-tools): Read, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Docker, NVIDIA Container Toolkit, and one visible GPU.

    From compatibility in the SKILL.md frontmatter.

Context cost

VLM BCQ Gap Analysis loads about 1.3k tokens when it runs, and up to ~1.5k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 523 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 523 words, ~1,315 tokens.

Download SKILL.mdSave it as .claude/skills/tao-analyze-gaps-vlm-bcq/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
tao-analyze-gaps-vlm-bcq
description
Extract false-positive and false-negative gaps from VLM binary-classification-question (BCQ, yes/no) predictions. Use when the user asks to "analyze VLM BCQ gaps", "extract VLM false positives and false negatives", or identify failure cases from a predictions JSON for DEFT root-cause analysis on a binary-classification VLM workflow.
allowed-tools
Read, Bash
compatibility
Requires Docker, NVIDIA Container Toolkit, and one visible GPU.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
gap-analysis, rcca, vlm, evaluation, false-positive, false-negative

VLM Binary Classification Gap Analysis

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Reads a VLM predictions JSON, compares each model response against ground truth, and writes FP/FN failure cases to a JSONL file with a summary report. Run it with a TAO Data Services spec file; the data-services entrypoint requires -e <spec>.

Purpose

After running a VLM on a binary yes/no evaluation task, the predictions need to be compared against ground truth to identify failure cases. This skill produces a structured list of FP (false positive) and FN (false negative) samples that downstream RCCA stages (e.g., cosmos generation, root cause analysis) consume to drive a DEFT iteration.

Usage

Generate a vlm_bcq_spec.yaml with the bundled helper:

bash
python3 skills/data/tao-analyze-gaps-vlm-bcq/scripts/prepare_vlm_bcq_spec.py \
  --predictions-json /path/to/results.json \
  --videos-dir /path/to/videos/root \
  --results-dir /path/to/output/gaps \
  --output-spec /path/to/output/gaps/vlm_bcq_spec.yaml

Omit --videos-dir when prediction video_id values are already absolute. The generated spec has this shape:

yaml
predictions_json: /path/to/results.json
videos_dir: ""
results_dir: /path/to/output/gaps

Set videos_dir when video_id values in the predictions are relative paths:

yaml
predictions_json: /path/to/results.json
videos_dir: /path/to/videos/root
results_dir: /path/to/output/gaps

Invoke the vlm_bcq action inside the TAO Toolkit data services container with -e <spec>:

bash
gap_analysis vlm_bcq -e /path/to/vlm_bcq_spec.yaml

Request exactly one GPU from the selected platform (compute_shape.gpus: 1, compute_shape.nodes: 1). VLM BCQ gap analysis does not perform GPU compute, but the Data Services image always calls nvidia-smi and fails when no GPU is visible. One is a GPU count, not a device ID; the platform selects the device.

After the run, surface the FP/FN counts from kpi_gaps_report.txt and point downstream stages at kpi_gaps.jsonl.

Inputs

  • config spec: YAML file passed with -e. Template: assets/default_vlm_bcq.yaml.
  • predictions_json: Path to predictions JSON file. Must be a JSON array where each item has video_id, response, and gt fields. response and gt are parsed with word-boundary matching — 'yes' or 'no' anywhere in the string is recognized. Samples where both or neither are present are skipped with a warning.
  • videos_dir (optional): Base directory for resolving relative video_id paths. If omitted, video_id values are used as absolute paths.
  • results_dir: Output directory for gap-analysis artifacts.

Predictions JSON format:

json
[
  {
    "video_id": "/path/to/video.mp4",
    "response": "Yes, there is a collision.",
    "gt": "B. No",
    "question": "Is there a collision?"
  }
]
Show full SKILL.md (200 more words)Show less

Outputs

  • kpi_gaps.jsonl: One JSON object per line for each FP/FN case. Fields: video_id (absolute path), error_type (FP or FN), question, ground_truth, response.
  • kpi_gaps_report.txt: Human-readable table with total FP/FN counts.

If no gaps are found, no files are written and a message is logged.

Spec Fields

ParameterRequiredDescription
predictions_jsonYesPath to predictions JSON file
results_dirYesOutput directory; created if it does not exist
videos_dirNoBase directory for resolving relative video_id paths

Keep the spec file and every path it references under the bind-mounted workspace so they resolve inside the container. Pass -e <spec> even if you also add Hydra overrides; current TAO Data Services entrypoints hard-require an experiment spec file before processing overrides.

Error Patterns

ErrorCauseFix
FileNotFoundErrorpredictions_json does not existCheck the path
requires the following argument: -e/--experiment_spec_fileThe container was launched without a spec fileWrite vlm_bcq_spec.yaml and pass gap_analysis vlm_bcq -e <spec>
ValueError: must be a JSON arrayPredictions file is not a listWrap predictions in [...]
ValueError: missing 'gt'/'response'/'video_id'A prediction item is missing a required fieldInspect and fix the predictions JSON
Samples silently skippedresponse or gt contains both or neither 'yes'/'no'Check logs for warnings; inspect those samples

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references, assets) in skills/tao-analyze-gaps-vlm-bcq of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • assets/default_vlm_bcq.yaml
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/skill_info.yaml
  • scripts/prepare_vlm_bcq_spec.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

VLM BCQ Gap Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

VLM BCQ Gap Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
VLM BCQ Gap Analysis this skillNVIDIA/skills3.6k—~1.3kAutomated safety check: NotesApache-2.0
Nemo Evaluator SDKOrchestra-Research/AI-Research-SKILLs13k2 repos~3.1kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0
Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassCustom licence
Yolo Detection 2026SharpAI/DeepCamera3.1k—~1.5kAutomated safety check: PassMIT
Code Model Evaluation HarnessOrchestra-Research/AI-Research-SKILLs13k4 repos~2.9kAutomated safety check: PassMIT

Similar skills

  • Nemo Evaluator SDK

    Orchestra-Research/AI-Research-SKILLs

    Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

    13k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Matlab Use Visual Inspection

    matlab/matlab-agentic-toolkit

    Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

    1.1k GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Yolo Detection 2026

    SharpAI/DeepCamera

    YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.

    3.1k GitHub stars~1.5k tokensUpdated 24 days ago
    AI & LLM EngineeringAuto-check passed
  • Code Model Evaluation Harness

    Orchestra-Research/AI-Research-SKILLs

    Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.

    13k GitHub starsUsed in 4 repos~2.9k tokens
    AI & LLM EngineeringAuto-check passed
  • OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU)

    3.1k GitHub stars~1.3k tokensUpdated 24 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about VLM BCQ Gap Analysis

What does VLM BCQ Gap Analysis do?

Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report. This skill reads a predictions JSON from a vision-language model run on a binary yes/no question task, checks each response against ground truth and extracts the false positive and false negative samples.txt, which downstream root-cause stages (such as cosmos generation and root cause analysis) use to drive a DEFT iteration.

When should I use VLM BCQ Gap Analysis?

VLM BCQ Gap Analysis fits situations like: pulling failure cases out of a VLM yes/no evaluation; preparing inputs for root-cause analysis in a DEFT iteration; getting false positive and false negative counts from a predictions JSON.

How do I install VLM BCQ Gap Analysis in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a claude-code`. Or copy the skill folder (skills/tao-analyze-gaps-vlm-bcq in NVIDIA/skills) into .claude/skills/tao-analyze-gaps-vlm-bcq in your project. Claude Code loads it when a task matches its description.

How do I install VLM BCQ Gap Analysis in Codex?

Run `npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a codex`. Or copy the skill folder (skills/tao-analyze-gaps-vlm-bcq in NVIDIA/skills) into .agents/skills/tao-analyze-gaps-vlm-bcq in your project. Codex loads it when a task matches its description.

Can I use VLM BCQ Gap Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-analyze-gaps-vlm-bcq, .gemini/skills/tao-analyze-gaps-vlm-bcq, .github/skills/tao-analyze-gaps-vlm-bcq and .opencode/skills/tao-analyze-gaps-vlm-bcq in your project.

What does VLM BCQ Gap Analysis need to run?

Going by SKILL.md and its folder, VLM BCQ Gap Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Docker, NVIDIA Container Toolkit and one visible GPU; A predictions JSON from a VLM binary-classification run; The TAO Toolkit data services container. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires Docker, NVIDIA Container Toolkit, and one visible GPU..

Does VLM BCQ Gap Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is VLM BCQ Gap Analysis safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does VLM BCQ Gap Analysis use?

VLM BCQ Gap Analysis is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does VLM BCQ Gap Analysis use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 218 tokens, read only when the agent opens those files.

What are the alternatives to VLM BCQ Gap Analysis?

Skills that share tags, products or a category with VLM BCQ Gap Analysis: Nemo Evaluator SDK (Orchestra-Research/AI-Research-SKILLs, 13k stars), Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 146 stars), Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars) and Yolo Detection 2026 (SharpAI/DeepCamera, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains VLM BCQ Gap Analysis?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.