Nemo Evaluator SDK
Orchestra-Research/AI-Research-SKILLs
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.
Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.
$ npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-analyze-gaps-vlm-bcq --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-analyze-gaps-vlm-bcq .claude/skills/tao-analyze-gaps-vlm-bcq && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-analyze-gaps-vlm-bcq" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-vlm-bcq into .claude/skills/tao-analyze-gaps-vlm-bcq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-gaps-vlm-bcq", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-vlm-bcqType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-analyze-gaps-vlm-bcq --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-analyze-gaps-vlm-bcq .agents/skills/tao-analyze-gaps-vlm-bcq && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-analyze-gaps-vlm-bcq" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-vlm-bcq into .agents/skills/tao-analyze-gaps-vlm-bcq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-gaps-vlm-bcq", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-analyze-gaps-vlm-bcq --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-analyze-gaps-vlm-bcq .cursor/skills/tao-analyze-gaps-vlm-bcq && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-analyze-gaps-vlm-bcq" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-vlm-bcq into .cursor/skills/tao-analyze-gaps-vlm-bcq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-gaps-vlm-bcq", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-analyze-gaps-vlm-bcq--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-analyze-gaps-vlm-bcq --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-analyze-gaps-vlm-bcq .gemini/skills/tao-analyze-gaps-vlm-bcq && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-analyze-gaps-vlm-bcq" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-vlm-bcq into .gemini/skills/tao-analyze-gaps-vlm-bcq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-gaps-vlm-bcq", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-analyze-gaps-vlm-bcqInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-analyze-gaps-vlm-bcq .github/skills/tao-analyze-gaps-vlm-bcq && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-analyze-gaps-vlm-bcq" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-vlm-bcq into .github/skills/tao-analyze-gaps-vlm-bcq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-gaps-vlm-bcq", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-analyze-gaps-vlm-bcq --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-analyze-gaps-vlm-bcq .opencode/skills/tao-analyze-gaps-vlm-bcq && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-analyze-gaps-vlm-bcq" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-analyze-gaps-vlm-bcq into .opencode/skills/tao-analyze-gaps-vlm-bcq/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-analyze-gaps-vlm-bcq", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-analyze-gaps-vlm-bcqCompares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report.
This skill reads a predictions JSON from a vision-language model run on a binary yes/no question task, checks each response against ground truth and extracts the false positive and false negative samples. Results go to kpi_gaps.jsonl with counts in kpi_gaps_report.txt, which downstream root-cause stages (such as cosmos generation and root cause analysis) use to drive a DEFT iteration.
A bundled helper, prepare_vlm_bcq_spec.py, builds the vlm_bcq_spec.yaml with the predictions file, an optional videos directory for relative video IDs and a results folder. The vlm_bcq action then runs inside the TAO Toolkit data services container with the spec passed through -e. The platform request must ask for exactly one GPU, because the image calls nvidia-smi even though the analysis itself does no GPU compute. If the session was not set up by the TAO skill bank plugin, tao-setup runs first.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Docker, NVIDIA Container Toolkit, and one visible GPU.
From compatibility in the SKILL.md frontmatter.
VLM BCQ Gap Analysis loads about 1.3k tokens when it runs, and up to ~1.5k if it reads all its reference files. Until then it costs about 90 tokens; SKILL.md has 523 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 523 words, ~1,315 tokens.
.claude/skills/tao-analyze-gaps-vlm-bcq/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Reads a VLM predictions JSON, compares each model response against ground truth, and writes FP/FN failure cases to a JSONL file with a summary report. Run it with a TAO Data Services spec file; the data-services entrypoint requires -e <spec>.
After running a VLM on a binary yes/no evaluation task, the predictions need to be compared against ground truth to identify failure cases. This skill produces a structured list of FP (false positive) and FN (false negative) samples that downstream RCCA stages (e.g., cosmos generation, root cause analysis) consume to drive a DEFT iteration.
Generate a vlm_bcq_spec.yaml with the bundled helper:
python3 skills/data/tao-analyze-gaps-vlm-bcq/scripts/prepare_vlm_bcq_spec.py \
--predictions-json /path/to/results.json \
--videos-dir /path/to/videos/root \
--results-dir /path/to/output/gaps \
--output-spec /path/to/output/gaps/vlm_bcq_spec.yamlOmit --videos-dir when prediction video_id values are already absolute. The generated spec has this shape:
predictions_json: /path/to/results.json
videos_dir: ""
results_dir: /path/to/output/gapsSet videos_dir when video_id values in the predictions are relative paths:
predictions_json: /path/to/results.json
videos_dir: /path/to/videos/root
results_dir: /path/to/output/gapsInvoke the vlm_bcq action inside the TAO Toolkit data services container with -e <spec>:
gap_analysis vlm_bcq -e /path/to/vlm_bcq_spec.yamlRequest exactly one GPU from the selected platform (compute_shape.gpus: 1,
compute_shape.nodes: 1). VLM BCQ gap analysis does not perform GPU compute,
but the Data Services image always calls nvidia-smi and fails when no GPU is
visible. One is a GPU count, not a device ID; the platform selects the device.
After the run, surface the FP/FN counts from kpi_gaps_report.txt and point downstream stages at kpi_gaps.jsonl.
-e. Template: assets/default_vlm_bcq.yaml.video_id, response, and gt fields. response and gt are parsed with word-boundary matching — 'yes' or 'no' anywhere in the string is recognized. Samples where both or neither are present are skipped with a warning.video_id paths. If omitted, video_id values are used as absolute paths.Predictions JSON format:
[
{
"video_id": "/path/to/video.mp4",
"response": "Yes, there is a collision.",
"gt": "B. No",
"question": "Is there a collision?"
}
]video_id (absolute path), error_type (FP or FN), question, ground_truth, response.If no gaps are found, no files are written and a message is logged.
| Parameter | Required | Description |
|---|---|---|
| predictions_json | Yes | Path to predictions JSON file |
| results_dir | Yes | Output directory; created if it does not exist |
| videos_dir | No | Base directory for resolving relative video_id paths |
Keep the spec file and every path it references under the bind-mounted workspace so they resolve inside the container. Pass -e <spec> even if you also add Hydra overrides; current TAO Data Services entrypoints hard-require an experiment spec file before processing overrides.
| Error | Cause | Fix |
|---|---|---|
FileNotFoundError | predictions_json does not exist | Check the path |
requires the following argument: -e/--experiment_spec_file | The container was launched without a spec file | Write vlm_bcq_spec.yaml and pass gap_analysis vlm_bcq -e <spec> |
ValueError: must be a JSON array | Predictions file is not a list | Wrap predictions in [...] |
ValueError: missing 'gt'/'response'/'video_id' | A prediction item is missing a required field | Inspect and fix the predictions JSON |
| Samples silently skipped | response or gt contains both or neither 'yes'/'no' | Check logs for warnings; inspect those samples |
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts, references, assets) in skills/tao-analyze-gaps-vlm-bcq of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
VLM BCQ Gap Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| VLM BCQ Gap Analysis this skillNVIDIA/skills | 3.6k | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | |
| Nemo Evaluator SDKOrchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | |
| Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit | 1.1k | — | ~3.1k | Automated safety check: Pass | Custom licence | |
| Yolo Detection 2026SharpAI/DeepCamera | 3.1k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Code Model Evaluation HarnessOrchestra-Research/AI-Research-SKILLs | 13k | 4 repos | ~2.9k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
matlab/matlab-agentic-toolkit
Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.
SharpAI/DeepCamera
YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.
Orchestra-Research/AI-Research-SKILLs
Benchmarks code generation models with the BigCode Evaluation Harness across HumanEval, MBPP, MultiPL-E and other suites using pass@k metrics.
SharpAI/DeepCamera
OpenVINO — real-time object detection via Docker (NCS2, Intel GPU, CPU)
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Compares a vision-language model's yes/no predictions with ground truth and writes the false-positive and false-negative cases to a JSONL file with a summary report. This skill reads a predictions JSON from a vision-language model run on a binary yes/no question task, checks each response against ground truth and extracts the false positive and false negative samples.txt, which downstream root-cause stages (such as cosmos generation and root cause analysis) use to drive a DEFT iteration.
VLM BCQ Gap Analysis fits situations like: pulling failure cases out of a VLM yes/no evaluation; preparing inputs for root-cause analysis in a DEFT iteration; getting false positive and false negative counts from a predictions JSON.
Run `npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a claude-code`. Or copy the skill folder (skills/tao-analyze-gaps-vlm-bcq in NVIDIA/skills) into .claude/skills/tao-analyze-gaps-vlm-bcq in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a codex`. Or copy the skill folder (skills/tao-analyze-gaps-vlm-bcq in NVIDIA/skills) into .agents/skills/tao-analyze-gaps-vlm-bcq in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-analyze-gaps-vlm-bcq -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-analyze-gaps-vlm-bcq, .gemini/skills/tao-analyze-gaps-vlm-bcq, .github/skills/tao-analyze-gaps-vlm-bcq and .opencode/skills/tao-analyze-gaps-vlm-bcq in your project.
Going by SKILL.md and its folder, VLM BCQ Gap Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Docker, NVIDIA Container Toolkit and one visible GPU; A predictions JSON from a VLM binary-classification run; The TAO Toolkit data services container. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires Docker, NVIDIA Container Toolkit, and one visible GPU..
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
VLM BCQ Gap Analysis is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 218 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with VLM BCQ Gap Analysis: Nemo Evaluator SDK (Orchestra-Research/AI-Research-SKILLs, 13k stars), Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 146 stars), Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars) and Yolo Detection 2026 (SharpAI/DeepCamera, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.