Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Vss Evaluate Caption Accuracy

skills CLI
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-evaluate-caption-accuracy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/vss-evaluate-caption-accuracy .claude/skills/vss-evaluate-caption-accuracy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vss-evaluate-caption-accuracy
GitHub stars
1.9k
Token cost
~2.1k tokens
SKILL.md length
865 words
Files
9 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

  • Changing frame selection
  • SKILL.md covers Purpose, When to Use, Prerequisites and Stages, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls bash, docker and python3; needs OPENAI_API_KEY and ANTHROPIC_API_KEY
  • Model settings and you need evidence there is no accuracy regression

What it does

Vss Evaluate Caption Accuracy is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and emit an accuracy and processing-time table. Use when changing frame selection, decode, or model settings and you need evidence there is no accuracy regression.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `evals/evals.json`, `scripts/aggregate_table.py` and `scripts/multi_judge.py`).

It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.

When your agent uses it

  • Changing frame selection
  • Model settings and you need evidence there is no accuracy regression

Example prompts

  • “/vss-evaluate-caption-accuracy”

Requirements

  • Python 3
  • A Bash shell
  • Docker
  • A credential in OPENAI_API_KEY
  • A credential in ANTHROPIC_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 3a75661. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • docker
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vss Evaluate Caption Accuracy loads about 2.1k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 865 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:43
    ts | Set `MODEL_PATH` in the deployment `.env`. Runs here used `Qwen3-VL-32B-Instruct`; any vLLM-compatible VLM works |
  • NoteMentions a .env fileSKILL.md:44
    SE=vllm-compatible` | In the deployment `.env` |
  • NoteMentions a .env fileSKILL.md:46
    | `OPENAI_API_KEY` | In the deployment `.env`. Only needed for the `gt` stage (ground truth is gpt-4.1) |
  • NoteMentions a .env fileSKILL.md:109
    rms. Unset inherits the container's own `.env`. When set, it is also used as `RTVI_MODEL_PATH_ALLOWLIST`, which `VLM_TRU

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit 3a75661, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 865 words, ~2,103 tokens.

Download SKILL.mdSave it as .claude/skills/vss-evaluate-caption-accuracy/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
vss-evaluate-caption-accuracy
description
Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and emit an accuracy and processing-time table. Use when changing frame selection, decode, or model settings and you need evidence there is no accuracy regression.
license
Apache-2.0
metadata.version
3.3.0-rc0
metadata.requires-vss
>=3.2.0,<4.0.0
metadata.author
NVIDIA Video Search and Summarization Team
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
metadata.tags
nvidia blueprint rt-vlm captions accuracy evaluation frame-selection

Purpose

Answer one question with evidence: did this RT-VLM change make captions worse?

A configuration change that saves processing time is only useful if caption quality holds. This skill captures captions twice over the same videos — once with the change (HYP) and once without (REF) — scores both against a ground truth using an LLM judge, and reports accuracy delta alongside time saved.

When to Use

  • Before/after changing RT-VLM frame selection, decode, or model settings, to prove there is no caption-quality regression
  • Quantifying accuracy delta and processing-time saved for a candidate config change
  • Verifying a change was behavior-neutral via frame-level provenance, not just matching totals

Not for deploying RT-VLM, general LVS/RTVI benchmarking (see vss-benchmark-video-summarization), or one-off caption spot checks with no paired baseline.

Prerequisites

Everything except the judge runs inside the RT-VLM container. The container needs a GPU, the model on disk, and the nvdsframeselector DeepStream plugin (shipped with DeepStream in the RT-VLM image).

RequirementNotes
Running RT-VLM containerDefault name rtvi_vlm-$USER. Start it before running any stage
Model weightsSet MODEL_PATH in the deployment .env. Runs here used Qwen3-VL-32B-Instruct; any vLLM-compatible VLM works
VLM_MODEL_TO_USE=vllm-compatibleIn the deployment .env
Source videosA directory of .mp4 files. Point DEDUP_DIR at it
OPENAI_API_KEYIn the deployment .env. Only needed for the gt stage (ground truth is gpt-4.1)
claude CLI on the hostThe judge shells out to it. It is not installed in the container
Deployment inside requires-vss>=3.2.0,<4.0.0, from this skill's front matter. Not enforced automatically — this skill has no preflight.sh, and it drives RT-VLM's OpenAI-compatible HTTP surface directly rather than a VSS deployment origin, so there is no origin here to check. Verify by hand if a deployment's version is in doubt: python3 <repo>/services/agent/scripts/check_vss_version.py <deployment-origin> --skill <this SKILL.md> (0 = compatible, 3 = incompatible, 1 = indeterminate)
Video paths and the scene map

scripts/run_captioning.py maps a filename to a short scene name in its SCENES dict. Add your own videos there:

python
SCENES = {
    "warehouse.mp4":               "warehouse",
    "GoPro5_10min_compressed.mp4": "new_warehouse",
    # "<your-file>.mp4":           "<scene-name>",
}

Scene names are what you pass on the command line; files are resolved inside DEDUP_DIR. A scene whose file is missing is skipped with a warning.

Model configuration

The skill does not choose a model — it inherits whatever the container is configured with, and sets only the knobs under test. Both arms run the same model so the comparison isolates the configuration change, not the checkpoint.

Stages

They run in different places, so invoke them separately:

StageWhereWhat
gtcontainerGround truth from gpt-4.1. Expensive — run once per video set and reuse via GT_SRC
capturecontainerPaired REF + HYP per scene, back to back
judgehostclaude-opus-4-8 scores REF and HYP against GT, per chunk
tableeitherAccuracy + time-saved markdown, optionally against a baseline run
bash
# 0. ground truth (once per video set)
docker exec -e DESC=my-run -w /workspace rtvi_vlm-$USER \
  bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh gt scene-a scene-b

# 1. capture — paired REF + HYP
docker exec -e DESC=my-run -w /workspace rtvi_vlm-$USER \
  bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh capture scene-a scene-b

# 2. judge — on the host, not in the container
DESC=my-run bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh judge scene-a scene-b

# 3. table
DESC=my-run bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh table scene-a scene-b

Replace -w /workspace with the path the repository is mounted at in your container.

Knobs

VarDefaultMeaning
DESCeval-<date>Run name. Everything lands in results/<DESC>/
VSS_REPO_DIR/workspacePath the repository is mounted at inside the container
RTVI_CONTAINERrtvi_vlm-$USERContainer name to exec into
MODEL_PATHunsetPin a checkpoint for both arms. Unset inherits the container's own .env. When set, it is also used as RTVI_MODEL_PATH_ALLOWLIST, which VLM_TRUST_REMOTE_CODE=true requires
DEDUP_DIR<repo>/videosDirectory holding the source videos
SFCunsetNVDS_FSELECT_STATIC_FRAME_COUNT — frames emitted for a chunk classified STATIC. Unset leaves the plugin default
GT_SRC$DESCRun to copy gt.txt from, so one ground truth serves many runs
BASELINEnoneRun to diff against in table
RESULTS_ROOT<skill>/resultsWhere run folders live
HYP_VERv1Which hyp_<desc>_vN.txt to judge
MAX_WORKERS32Judge concurrency
Show full SKILL.md (293 more words)Show less

Why REF and HYP are captured paired

REF caption generation is nondeterministic run to run. Judging HYP against a REF captured in a different session moved per-scene deltas by up to 0.05 — larger than most effects being measured. capture therefore runs HYP and REF back to back per scene in one session. Only GT is reused, because it is expensive and comes from a different model.

Reading the output

table writes results/<DESC>/summary.md: accuracy and processing time per scene, LLM-judge entity/event F1 detail, and — with BASELINE — the incremental effect versus that run. Totals are chunk-weighted, so a 60-chunk scene does not carry the same weight as a 14-chunk one.

Two cautions:

  • Check the noise floor first. Repeat runs of an identical configuration have spanned ~0.01 on a single scene, and the REF/HYP delta has changed sign between them. A delta smaller than that is "unchanged", not "improved". If a result matters, repeat it.
  • A flat combined score can hide offsetting axes. combined_score_macro_0_1 averages entity, event, critical-event and interaction F1. Interaction F1 is often a handful of samples and carries little signal, so it can mask a real entity-F1 move. Read the F1 detail block, not just the combined column.

Verifying that a change was behaviour-neutral

Frame-level provenance is stronger evidence than matching totals. When frame selection is active the plugin logs one line per chunk:

bash
grep -oE 'EOS OF-only -> [0-9]+' results/<DESC>/server_logs/hyp_<scene>.log \
  | grep -oE '[0-9]+$' | tail -n <chunks> | sort -n | uniq -c

Identical per-chunk frame counts across two runs mean the selector chose the same frames. Matching chunk counts alone do not. The first values belong to pipeline warmup rather than the scene, so take the trailing <chunks> entries.

Contents

  • scripts/run_eval.sh — the four stages
  • scripts/run_captioning.py — capture engine (gt / ref / hyp)
  • scripts/multi_judge.py — judge engine; uses the local claude CLI, so no ANTHROPIC_API_KEY is needed
  • scripts/score.py — per-chunk scoring into summary.csv
  • scripts/aggregate_table.py — the report

© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts) in skills/benchmarking/vss-evaluate-caption-accuracy of NVIDIA-AI-Blueprints/video-search-and-summarization.

  • SKILL.md
  • .gitignore
  • evals/evals.json
  • scripts/aggregate_table.py
  • scripts/multi_judge.py
  • scripts/run_captioning.py
  • scripts/run_eval.sh
  • scripts/score.py
  • skill-card.md

Open the folder on GitHubat commit 3a75661

Compare with similar skills

Vss Evaluate Caption Accuracy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vss Evaluate Caption Accuracy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vss Evaluate Caption Accuracy this skillNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~2.1kAutomated safety check: NotesApache-2.0
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Looperksimback/looper710—~2.7kAutomated safety check: NotesMIT
Agent Eval Engineeringlangchain-ai/langchain-skills1.3k—~4kAutomated safety check: PassMIT
Quality FlywheelGoogleCloudPlatform/vertex-ai-samples792—~2kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Looper

    ksimback/looper

    Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.

    710 GitHub stars~2.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Agent Eval Engineering

    langchain-ai/langchain-skills

    Official

    Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.

    1.3k GitHub stars~4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Quality Flywheel

    GoogleCloudPlatform/vertex-ai-samples

    Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.

    792 GitHub stars~2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    cloudnative-co/claude-code-starter-kit

    Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.

    153 GitHub starsUsed in 9 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA-AI-Blueprints/video-search-and-summarization

All 22 skills in this repo
  • Benchmark Video Search

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

    1.9k GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video or ingest or delete a source for search.

    1.9k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Rtvi Vlm Perf Testing

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.

    1.9k GitHub stars~8.6k tokensUpdated today
    Auto-check: notes
  • Vss Build Vision AI

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…

    1.9k GitHub stars~15k tokensUpdated today
    Auto-check: notes
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Vss Manage Alerts

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when operating VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, verification verdicts, on-demand verification, always-on operation, Slack…

    1.9k GitHub stars~11k tokensUpdated today
    Auto-check: warnings

Questions about Vss Evaluate Caption Accuracy

What does Vss Evaluate Caption Accuracy do?

Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…. Vss Evaluate Caption Accuracy is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and emit an accuracy and processing-time table.

When should I use Vss Evaluate Caption Accuracy?

Vss Evaluate Caption Accuracy fits situations like: changing frame selection; model settings and you need evidence there is no accuracy regression.

How do I install Vss Evaluate Caption Accuracy in Claude Code?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a claude-code`. Or copy the skill folder (skills/benchmarking/vss-evaluate-caption-accuracy in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/vss-evaluate-caption-accuracy in your project. Claude Code loads it when a task matches its description.

How do I install Vss Evaluate Caption Accuracy in Codex?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a codex`. Or copy the skill folder (skills/benchmarking/vss-evaluate-caption-accuracy in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/vss-evaluate-caption-accuracy in your project. Codex loads it when a task matches its description.

Can I use Vss Evaluate Caption Accuracy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vss-evaluate-caption-accuracy, .gemini/skills/vss-evaluate-caption-accuracy, .github/skills/vss-evaluate-caption-accuracy and .opencode/skills/vss-evaluate-caption-accuracy in your project.

What does Vss Evaluate Caption Accuracy need to run?

Going by SKILL.md and its folder, Vss Evaluate Caption Accuracy needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (bash, docker and python3) and credentials named OPENAI_API_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; A Bash shell; Docker; A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY.

Does Vss Evaluate Caption Accuracy access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vss Evaluate Caption Accuracy safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Vss Evaluate Caption Accuracy use?

Vss Evaluate Caption Accuracy is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vss Evaluate Caption Accuracy use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vss Evaluate Caption Accuracy?

Skills that share tags, products or a category with Vss Evaluate Caption Accuracy: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vss Evaluate Caption Accuracy?

NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,917 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.