Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Vss Benchmark Vlm QA

skills CLI
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-benchmark-vlm-qa -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-benchmark-vlm-qa --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/vss-benchmark-vlm-qa .claude/skills/vss-benchmark-vlm-qa && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vss-benchmark-vlm-qa
GitHub stars
1.9k
Token cost
~2.5k tokens
SKILL.md length
1,118 words
Files
5 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset.

  • Tasks that involve Summarization
  • SKILL.md covers When to use, When not to use, Prerequisites and Run, plus 2 more sections
  • Runs Python and Shell scripts from its folder; calls uv, python3 and docker; reaches artifactory.pdx.nvidia.com; needs NGC_API_KEY and NVDATASET_API_KEY
  • Tasks that involve Structured output and tool calling

What it does

Vss Benchmark Vlm QA is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. Replaces the deprecated nat eval / vss-agent QA path. Not for tool-calling or trajectory evaluation, and not for LVS summarization throughput.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts (for example `scripts/benchmark_vlm_qa.py`, `scripts/run_vlm_qa_benchmark.sh` and `scripts/tests/conftest.py`).

It sits in AI & LLM Engineering, covering Summarization, Structured output and tool calling and LLM evaluation. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Summarization
  • Tasks that involve Structured output and tool calling
  • Tasks that involve LLM evaluation

Example prompts

  • “/vss-benchmark-vlm-qa”

Requirements

  • Python 3
  • A Bash shell
  • Docker
  • A credential in NVDATASET_API_KEY
  • A credential in NVDATASET_NGC_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • python3
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • artifactory.pdx.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NGC_API_KEY
    • NVDATASET_API_KEY
    • NVDATASET_NGC_API_KEY
    • EVAL_LLM_JUDGE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vss Benchmark Vlm QA loads about 2.5k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 1,118 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 1,118 words, ~2,536 tokens.

Download SKILL.mdSave it as .claude/skills/vss-benchmark-vlm-qa/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
vss-benchmark-vlm-qa
description
Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. Replaces the deprecated nat eval / vss-agent QA path. Not for tool-calling or trajectory evaluation, and not for LVS summarization throughput.
license
Apache-2.0
metadata.version
3.3.0-rc0
metadata.requires-vss
>=3.3.0,<4.0.0
metadata.author
NVIDIA Video Search and Summarization Team
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
metadata.tags
nvidia blueprint performance benchmarking vlm qa

Benchmark video Q&A via vss vlm

Measure accuracy (LLM-as-judge vs ground truth) and latency of end-to-end video question answering by calling vss vlm run against a deployed Cosmos Reason 3 RT-VLM. Questions and clips come from DSS dataset vss-devx-base (nvdataset).

This replaces docker exec vss-agent nat eval for the QA slice. It does not score tool-calling or trajectories.

When to use

  • The user asks to benchmark / evaluate VLM video Q&A after vss-agent / NAT eval was removed.
  • The user wants latency and answer accuracy on vss-devx-base.

When not to use

  • Tool-calling or trajectory evaluation — out of scope.
  • LVS summarization throughput — use vss-benchmark-video-summarization.
  • Ad-hoc single questions — use /vss-ask-video.

Prerequisites

  • A VSS stack with RT-VLM serving Cosmos Reason 3, and vss configure already run so vss configure check lists rt_vlm as ok and vst as ok.

    Configure with a routable address, not localhost. Clips are addressed as VIOS sensors so RT-VLM fetches them by URL; the URL VIOS mints is built from the configured origin. A loopback origin mints a loopback URL, which means nothing inside the RT-VLM container, so the CLI falls back to inlining the clip as base64 and the VLM rejects anything large with HTTP 422 ... content ... valid string. vss configure --base-url http://<host-ip>:7777 avoids that — --base-url is a vss configure flag, not a benchmark one. --inline-media is a benchmark flag; it forces the old inline behaviour and is only safe for clips under ~10 MB.

  • uv and this checkout. benchmark_vlm_qa.py runs the CLI as uv run --project services/agent --no-dev --extra cli vss.

  • A deployment inside this skill's requires-vss range (>=3.3.0,<4.0.0 — the floor is the Cosmos Reason 3 stack the published baselines were measured on, not a CLI capability: the CLI comes from this checkout either way). This is not enforced automatically. Unlike vss-benchmark-video-summarization, this skill has no preflight.sh, and adding one just to carry a single check would be out of proportion; vss configure check is already the gate it runs. Verify by hand when in doubt:

    bash
    python3 <repo>/services/agent/scripts/check_vss_version.py \
      <deployment-origin> --skill <repo>/skills/benchmarking/vss-benchmark-vlm-qa/SKILL.md

    Exit 0 = compatible, 3 = incompatible, 1 = could not be determined.

  • The nvdataset CLI. It is not on PyPI, and the index used by the old deep-search eval (urm.nvidia.com/.../sw-ngc-data-platform-pypi) returns 403. Install from the documented read-only index instead — no credentials needed:

    bash
    uv tool install --index https://artifactory.pdx.nvidia.com/artifactory/api/pypi/sw-ngc-data-platform-pypi-local/simple nvdataset
  • DSS access, one of:

    • NVDATASET_API_KEY — a Personal Key from org.ngc.nvidia.com/setup/personal-keys scoped to the service NVIDIA Dataset Service, with the NGC org switched to the one owning the dataset. This is not the global NGC key used by the NGC CLI; a global key returns 403. NVDATASET_NGC_API_KEY and NGC_API_KEY are also read, in that order, for backward compatibility only — the run prints the variable it picked as dss credential: <name>, so check that line if a 403 surprises you.
    • nvdataset auth login (Starfleet SSO), which needs no key. Add --flow device on a remote box with no browser. Group access requires membership in ngc-datasetservice-viewer-<tenant>-<group> (reader) or ...-user-... (writer).

    Plus tenancy, which SSO does not supply — after auth login, nvdataset auth status still reports "tenant_id": null and every call fails with Did not find tenant_id. The script names no tenant, so set one yourself: export NVDATASET_TENANTID and NVDATASET_GROUPID, or save them once with nvdataset auth context add. Ask the dataset's owning team for its coordinates. Another dataset needs no change to the script.

  • An OpenAI-compatible judge LLM: EVAL_LLM_JUDGE_BASE_URL and EVAL_LLM_JUDGE_NAME, authenticated with EVAL_LLM_JUDGE_API_KEY. NGC_API_KEY is deliberately not sent to non-NVIDIA judge hosts — it is set for the dataset download and must not reach a third party. Any chat-completions endpoint will do; the judge moves absolute scores on its own, so hold it fixed across runs you mean to compare, and read judge_model in summary.json before comparing two numbers. --skip-judge gives latency only.

Bootstrap is in the repo-root AGENTS.md. Do not construct RT-VLM URLs; vss vlm run reads the recorded config.

Show full SKILL.md (493 more words)Show less

Run

bash
export NVDATASET_API_KEY=<personal-key>            # or: nvdataset auth login [--flow device]
export NVDATASET_TENANTID=<tenant>                 # SSO does not set this; see Prerequisites
export NVDATASET_GROUPID=<group>
export EVAL_LLM_JUDGE_BASE_URL="${LLM_BASE_URL}"   # OpenAI-compat origin, e.g. http://127.0.0.1:8000
export EVAL_LLM_JUDGE_NAME="${LLM_NAME}"

# Optional: already-extracted dataset
# export VSS_EVAL_DATASET=/path/to/vss-devx-base

<repo>/skills/benchmarking/vss-benchmark-vlm-qa/scripts/run_vlm_qa_benchmark.sh \
  --dataset-name vss-devx-base \
  --dataset-file dataset_single_turn.json

Both dataset flags are required — the script carries no default dataset, so it never assumes one team's DSS coordinates.

Useful flags (forwarded to benchmark_vlm_qa.py):

FlagPurpose
--dry-runResolve QA items and video files; no VLM calls
--limit NFirst N QA items (smoke)
--skip-judgeLatency only
--skip-downloadUse an already-downloaded vss-devx-base
--timeout SECPassed through as vss vlm run --timeout (default 300)
--num-frames NFrame budget (default 20, matching the old RT-VLM agent config)
--model IDOverride the RT-VLM model vss configure recorded

Outputs under <dataset>/../../results/vlm_qa/ (or --output-dir):

  • summary.json — mean accuracy, latency mean / p50 / p90 / p95 / p99, and the model the deployment reported serving, so a number is never left unattributable
  • qa_evaluator_output.json — per-item judge scores (same shape as NAT QA output)
  • latency_summary.json — per-item wall-clock around vss vlm run
  • workflow_output.json — raw answers
  • summary.csv

Rules

  • Drive the VLM only through vss vlm run. Never POST /generate or hand-built /v1/chat/completions.
  • Do not wrap vss in retries. --timeout is the bound; the script adds only a hard kill 60 s past it, so a CLI that never returns cannot cost the whole run. A killed item is recorded as an error naming the watchdog, never as a low score.
  • Items must declare evaluation_method containing qa and carry a text ground_truth. Report, trajectory-only, and unmarked items are skipped.

Failures

Branch on the exit code; never scrape stdout for the word "error".

ExitMeaningWhat to do
0Every item answeredRead summary.json
2Precondition wrong — a dataset flag missing, no DSS credential, no judge configured, dataset or videos not found, no QA itemsFix the setup. Re-running unchanged fails identically
3The download failed, or at least one item erroredRead each item's error in summary.json

A vss call that exits 4 (service missing from the recorded config) surfaces as an item error, so the run ends at exit 3 — the fix is vss configure, not a flag.

Failures worth recognising by their message:

  • HTTP 422 ... content ... valid string on the big clips — the recorded origin is loopback, so clips are being inlined as base64. Reconfigure with a routable address.
  • Did not find tenant_id — SSO signed you in but selected no tenant. Export NVDATASET_TENANTID, or nvdataset auth context use.
  • LLM judge HTTP 403 ... key_model_access_denied or 400 Invalid model name on every item — the judge id is not what that gateway calls the model. Gateways that front several providers usually want a fully-qualified id and reject the bare name. GET <judge-base-url>/models lists the ids the key may use; copy one verbatim into EVAL_LLM_JUDGE_NAME. The VLM answers are unaffected, so only scoring is lost.
  • An item error naming the watchdog — the CLI never returned and was killed at --timeout + 60 s. That is recorded as an error, never as a low score. Do not retry.
  • Accuracy far from the ~0.465 baseline is not a harness failure. The judge model and --num-frames both move it; check judge_model and model_served before filing.

Implementation: scripts/benchmark_vlm_qa.py, tested by scripts/tests/. Dataset download contract: README_eval.md.

© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in skills/benchmarking/vss-benchmark-vlm-qa of NVIDIA-AI-Blueprints/video-search-and-summarization.

  • SKILL.md
  • scripts/benchmark_vlm_qa.py
  • scripts/run_vlm_qa_benchmark.sh
  • scripts/tests/conftest.py
  • scripts/tests/test_benchmark_vlm_qa.py

Open the folder on GitHubat commit fdb6a7a

Compare with similar skills

Vss Benchmark Vlm QA next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vss Benchmark Vlm QA compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vss Benchmark Vlm QA this skillNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.5kAutomated safety check: PassApache-2.0
Dspy Evaluation Harnessintertwine/dspy-agent-skills278—~1.5kAutomated safety check: PassMIT
Foundation Modelsjohnrogers/claude-swift-engineering231—~869Automated safety check: PassMIT
Building Agent Systemstelagod/code-abyss244—~691Automated safety check: PassMIT
MLtelagod/code-abyss244—~566Automated safety check: PassMIT
Contextpilot SavingsEfficientContext/ContextPilot141—~1.4kAutomated safety check: PassMIT

Similar skills

  • Dspy Evaluation Harness

    intertwine/dspy-agent-skills

    Build DSPy evaluation harnesses with rich-feedback metrics that are essential for GEPA optimization.

    278 GitHub stars~1.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Foundation Models

    johnrogers/claude-swift-engineering

    A skill your agent uses when implementing on-device AI with Apple's Foundation Models framework (iOS 26+), building summarization/extraction/classification features, or using @Generable for…

    231 GitHub stars~869 tokensUpdated 8 mo ago
    AI & LLM EngineeringAuto-check passed
  • Building Agent Systems

    telagod/code-abyss

    AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt…

    244 GitHub stars~691 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • ML

    telagod/code-abyss

    Machine learning and LLM engineering judgment, distilled from a stronger model - invoke when DECIDING whether/how to use ML or an LLM for a task (prompt vs RAG vs fine-tune vs classical); working…

    244 GitHub stars~566 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Contextpilot Savings

    EfficientContext/ContextPilot

    A skill your agent uses when a user asks how many tokens (or how much context/cost) ContextPilot has saved, or wants a ContextPilot savings status/summary inside Hermes Agent — e.g.

    141 GitHub stars~1.4k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Prompt Engineer

    Jeffallan/claude-skills

    Designs, tests and refines LLM prompts: zero-shot, few-shot and chain-of-thought patterns, system prompts, structured output schemas and evaluation test suites.

    12k GitHub stars~1.5k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA-AI-Blueprints/video-search-and-summarization

All 22 skills in this repo
  • Benchmark Video Search

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

    1.9k GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…

    1.9k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Rtvi Vlm Perf Testing

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.

    1.9k GitHub stars~8.6k tokensUpdated yesterday
    Auto-check: notes
  • Vss Build Vision AI

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…

    1.9k GitHub stars~15k tokensUpdated yesterday
    Auto-check: notes
  • Vss Evaluate Caption Accuracy

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

    1.9k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check: notes
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Vss Benchmark Vlm QA

What does Vss Benchmark Vlm QA do?

Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. Vss Benchmark Vlm QA is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset.

When should I use Vss Benchmark Vlm QA?

Vss Benchmark Vlm QA fits situations like: tasks that involve Summarization; tasks that involve Structured output and tool calling; tasks that involve LLM evaluation.

How do I install Vss Benchmark Vlm QA in Claude Code?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-benchmark-vlm-qa -a claude-code`. Or copy the skill folder (skills/benchmarking/vss-benchmark-vlm-qa in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/vss-benchmark-vlm-qa in your project. Claude Code loads it when a task matches its description.

How do I install Vss Benchmark Vlm QA in Codex?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-benchmark-vlm-qa -a codex`. Or copy the skill folder (skills/benchmarking/vss-benchmark-vlm-qa in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/vss-benchmark-vlm-qa in your project. Codex loads it when a task matches its description.

Can I use Vss Benchmark Vlm QA in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-benchmark-vlm-qa -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vss-benchmark-vlm-qa, .gemini/skills/vss-benchmark-vlm-qa, .github/skills/vss-benchmark-vlm-qa and .opencode/skills/vss-benchmark-vlm-qa in your project.

What does Vss Benchmark Vlm QA need to run?

Going by SKILL.md and its folder, Vss Benchmark Vlm QA needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (uv, python3 and docker) and credentials named NGC_API_KEY, NVDATASET_API_KEY, NVDATASET_NGC_API_KEY and EVAL_LLM_JUDGE_API_KEY. Our summary lists: Python 3; A Bash shell; Docker; A credential in NVDATASET_API_KEY; A credential in NVDATASET_NGC_API_KEY.

Does Vss Benchmark Vlm QA access the network?

SKILL.md names 1 domain. In commands or code: artifactory.pdx.nvidia.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Vss Benchmark Vlm QA safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Vss Benchmark Vlm QA use?

Vss Benchmark Vlm QA is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vss Benchmark Vlm QA use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vss Benchmark Vlm QA?

Skills that share tags, products or a category with Vss Benchmark Vlm QA: Dspy Evaluation Harness (intertwine/dspy-agent-skills, 278 stars), Foundation Models (johnrogers/claude-swift-engineering, 231 stars), Building Agent Systems (telagod/code-abyss, 244 stars) and ML (telagod/code-abyss, 244 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vss Benchmark Vlm QA?

NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.

Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.