Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Benchmark Video Search

skills CLI
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization benchmark-video-search --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/benchmark-video-search .claude/skills/benchmark-video-search && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-video-search
GitHub stars
1.9k
Token cost
~4.3k tokens
SKILL.md length
2,269 words
Files
21 (incl. scripts, references)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

  • Works in 6 steps: Configure the CLI → Dry run → Verify the deployment is actually indexed → …
  • The user asks to benchmark video search
  • SKILL.md covers Instructions, Purpose, When not to use this skill and Routing, plus 10 more sections
  • Runs Python scripts from its folder; calls uv, curl and python3

What it does

Benchmark Video Search is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths. Use when the user asks to "benchmark video search", "compare search recall between builds", or "profile search latency" on a deployed VSS profile.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including scripts and reference files (for example `evals/evals.json`, `references/dataset-format.md` and `references/flags.md`).

It sits in AI & LLM Engineering. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.

When your agent uses it

  • The user asks to benchmark video search
  • Compare search recall between builds
  • Profile search latency on a deployed VSS profile

Example prompts

  • “benchmark video search”
  • “compare search recall between builds”
  • “profile search latency”
  • “/benchmark-video-search”

Requirements

  • Python 3
  • Docker

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Configure the CLI
  2. Dry run
  3. Verify the deployment is actually indexed
  4. Ingest
  5. Run the benchmark
  6. Report

What it can do on your machine

Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 10 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • curl
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Video Search loads about 4.3k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 2,269 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 2,269 words, ~4,291 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-video-search/SKILL.md (or your agent's skills folder). This skill also uses 20 other files; get the full folder from GitHub.
name
benchmark-video-search
description
Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths. Use when the user asks to "benchmark video search", "compare search recall between builds", or "profile search latency" on a deployed VSS profile.
license
Apache-2.0
metadata.version
3.3.0-rc0
metadata.author
NVIDIA Video Search and Summarization Team
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
metadata.tags
nvidia blueprint search retrieval benchmarking evaluation

Instructions

Follow the routing table, then the numbered steps. Execute each step in order on a first run; for a repeat run against an already-indexed deployment go straight to Step 5. Reference material is in references/, the runner is in scripts/.

Purpose

Answers two questions about a search profile deployment: does retrieval find the right footage, and where does a query's time go. It drives the vss CLI — the same path the product uses, since the agent adapter shells out to vss search run <mode> --raw for every search — so the numbers describe what ships rather than a REST endpoint that is being retired.

When not to use this skill

  • For a single video-search question, use vss-search-archive.
  • For an end-to-end benchmark of the OpenClaw chat route, this runner does not measure that path; do not present CLI timings as chat timings.
  • For a historical benchmark of the retired agent REST search endpoint, use the external legacy run_eval.py noted in references/troubleshooting.md.

Routing

SituationAction
User asks to benchmark the OpenClaw chat routeThis runner does not measure it; explain the gap instead of presenting CLI timings as chat timings
User asks to benchmark the retired agent REST search endpointThis runner does not measure it; for historical comparison, use the external legacy run_eval.py noted in references/troubleshooting.md
No search profile deployed in this sessionUse vss-build-vision-ai to deploy the stock search profile with its in-stack agent REST API and in-stack LLM retained. Explicitly name both requirements in the build request so it skips the harness question (which would remove the agent and possibly the LLM); do not select a NemoClaw-only or CLI-only harness. Record the reachable agent endpoint and unified origin, then return here. Check agent /health and LLM /v1/models, run Step 3 to inspect indices, and complete Step 4 ingestion and its index probe before scoring
User did not give an endpointAsk for it. Do not guess, and do not default to localhost
Endpoint uses localhost or 127.0.0.1This is valid when the runner is on the deployment host. Verify the actual VST clip URL is reachable from RT-VLM before trusting critic-filtered metrics
User did not give a datasetAsk which dataset and where its --data-dir is. Do not invent one
Elasticsearch has no mdx-* indicesIngestion has not completed. Run Step 4; do not report zero scores as a quality result
User asks to reuse what is already ingestedAdd --skip-ingest. It cannot be combined with --clear or --only-dataset
User asks to analyse an existing result fileSkip to Step 6; read the JSON, run nothing
Every query returns 0 hitsStop. Diagnose with Step 3 before reporting anything — this is nearly always missing indices, not poor retrieval
Critic-filtered metrics are all NAThe critic never ran. Check the clip-URL prerequisite below before concluding anything about verification quality

The default run

Unless the user says otherwise, this is the run. Ask which dataset to use and, if missing, which deployment endpoint to target.

bash
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
    --data-dir /path/to/datasets --dataset DATASET \
    --skip-download --skip-existing --name RUN_NAME

Five defaults, each load-bearing:

  • --skip-existing is non-destructive and avoids duplicate uploads on a shared deployment. --only-dataset and --clear are explicit destructive maintenance modes. Both require --confirm-delete in the same invocation; there is no interactive prompt. Use either only when the user requested deletion, and report the printed deletion inventory.

  • Live decomposition is on by default. The LLM origin is derived from --endpoint (same host, port 30081), so there is no flag to forget. Every query is decomposed the way the deployed agent decomposes it, and that choice picks the retrieval path. Override with --llm-url / --llm-port; turn it off with --no-decompose (replays stored routes) or --fixed-search-path (ignores stored routes for a single-path baseline).

    If the NIM is unreachable the run replays dataset routes when present; otherwise it falls back to embed, which needs nothing but the query text. The fallback records flow.query.live_decomposition.fell_back_to in the result file. Check that field and Search paths: before quoting numbers.

    --no-decompose also replays dataset routes when present; it does not force one path. Use --fixed-search-path only for a deliberately single-path baseline. Nothing else is inferred: a devset_provenance.json sidecar is never loaded automatically. To use it as an answer key, pass its full path with --decompositions; scoring against perfect routing is a different experiment from scoring the live decomposer.

  • The pre-decomposition question is sent with every query as --original-query. Retrieval ignores it; the critic verifies against it instead of a question the library rebuilds out of --query and --attribute. Without it the CLI's critic is asked a different question from the REST flow's, so the two flows' rejection rates are not comparable — and on the object path there is no question at all, so verification is skipped outright. Turn it off only with --no-original-query, and only to reproduce a result file captured before this existed.

    A vss older than the flag is detected by one --help probe at startup and the run continues without it, with a warning. Do not quote critic-filtered metrics from a run that printed that warning.

  • Ingest is vst-direct: upload to VIOS, let its webhooks drive perception. That is what the UI does — it stopped calling the agent's ingest API — and it is the only flow that also triggers RT-VLM tagging. The timestamp anchor rides in the upload metadata, not in anything /complete does, so scores stay comparable with agent-3step baselines.

    What agent-3step gave for free was proof: /complete returns chunks_processed, and a zero failed the upload. vst-direct has no such step, so a post-ingest index probe replaces it — one embed query, retried, before the real run starts. Registered in VST is not the same as indexed in Elasticsearch, and without this check an unindexed deployment scores 0.0 across the board and reads as a retrieval collapse.

    If the probe aborts a run, check the deployed VIOS notification config and that RTVI_EMBED_MODEL matches the webhook's model string — RT-Embed answers a mismatch with HTTP 200 and inference: false. The Docker and Helm search profiles enable webhooks; generic VIOS chart defaults may not. If the deployed webhooks are disabled, use --ingest-flow agent-3step. Never pass --skip-index-probe on a run whose numbers you intend to quote.

  • Concurrency stays at 1 (the script's default). Concurrent queries contend for the same VLM and embedding services, so per-stage latencies inflate and stop describing a single query.

Deviate only on request: --skip-ingest to reuse what is there, --clear to wipe the whole deployment, --subset to narrow the slice, --concurrency N to trade latency fidelity for wall-clock, --no-decompose to replay stored routes, or --fixed-search-path for a single-path baseline.

Never add either decomposition override to a run the user called a live-routing eval. Both measure a different flow; the result records the routing mode.

Prerequisites

RequirementHow to check
Search profile agent reachablecurl -sf --connect-timeout 5 --max-time 10 ${ENDPOINT}/health returns 200
Unified HAProxy origin routes the servicescurl -sf --connect-timeout 5 --max-time 10 ${ORIGIN}/vst/api/v1/sensor/version returns 200. ORIGIN is usually the endpoint host on port 7777; use it for vss configure and the Step 3 Elasticsearch probe. The runner separately derives VST's direct ingress on port 30888 (--vst-port default); do not pass the HAProxy port as --vst-port
vss CLI runs from this checkoutvss search run --help exits 0
CLI points at the right originvss configure show reports the ORIGIN above
Dataset present locallytest -f ${DATA_DIR}/${DATASET}/dataset.json — layout in references/dataset-format.md
Python 3.10+ for this runner; Python 3.13–3.14 for a separately installed vss CLIpython3 --version for the runner; use a Python 3.13 or 3.14 environment when installing libs/vss/core and libs/vss/cli (their requires-python is >=3.13,<3.15), then check vss search run --help there
LLM reachable (else the run falls back and says so)curl -sf --connect-timeout 5 --max-time 10 http://HOST:30081/v1/models returns 200
LLM model selected unambiguouslyIf /v1/models lists multiple IDs, pass --llm-model matching the deployed agent's model
Critic clip URL reachable from RT-VLMInspect a returned VST videoUrl and test that exact URL from the RT-VLM container. The CLI uses video_url_scope="internal"; CLI ORIGIN may legitimately be localhost on the host

Agent /health confirms the process is reachable, not that VIOS and the search indices are ready. Step 3 and the post-ingest index probe are the retrieval readiness gates.

Show full SKILL.md (920 more words)Show less

Step 1 — Configure the CLI

The CLI reads ~/.vss/config.json, which is separate state from --endpoint and points at a different port: it discovers services by path prefix on the unified origin, which the agent port does not route. Getting this wrong makes every query exit 4.

bash
vss configure --base-url http://HOST:7777
vss configure show

localhost is valid here when the runner runs on the deployment host. This URL controls where the CLI reaches the services; it does not determine the clip URL sent to RT-VLM. If verification fails, inspect the VST-returned videoUrl and test that exact address from the RT-VLM container.

Step 2 — Dry run

Reports what would happen and contacts nothing.

bash
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
    --data-dir /path/to/datasets --dataset DATASET --dry-run

Step 3 — Verify the deployment is actually indexed

Registration with VST is not ingestion. A deployment can list every sensor and still have an empty Elasticsearch, in which case every query returns zero hits and every metric reads 0.0000 — which looks like catastrophic retrieval quality and is not.

bash
curl -s --connect-timeout 5 --max-time 10 "http://HOST:7777/elasticsearch/_cat/indices?h=index,docs.count"

Expect mdx-embed-filtered-*, mdx-behavior-* and mdx-raw-* with non-zero counts. mdx-behavior-* is what the attribute and fusion paths read; if it is missing, only embed can score. If the indices are absent, go to Step 4.

Step 4 — Ingest

The default vst-direct flow uploads to VIOS and lets its webhooks trigger perception. It does not call the agent's /complete endpoint or receive a chunks_processed count. The runner waits for VST registration, then probes the search index before scoring. Use --ingest-flow agent-3step only when the deployment cannot run the webhook flow; see references/troubleshooting.md.

bash
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
    --data-dir /path/to/datasets --dataset DATASET --skip-download --skip-existing

Expect this to be slow while the webhook pipeline indexes the videos:

  • A registered source is not necessarily indexed. The post-ingest index probe must pass before treating zero hits as a retrieval result.
  • If webhooks are disabled, select --ingest-flow agent-3step explicitly. That flow calls /complete, which can return a transient 502 and is retried.
  • --skip-existing is the default: ingest what is missing and delete nothing.
  • --only-dataset deletes foreign sources only when explicitly requested with --confirm-delete; report its candidate inventory before it runs.
  • --clear --confirm-delete deletes every source including other people's. Use it only on explicit request.
  • --clear and --only-dataset list sources via --vst-url but delete through the agent endpoint. If those URLs have different hosts, verify they belong to the same deployment before deletion; otherwise source IDs from one stack could be sent to another.

Then re-run Step 3. Indices are lazy; they appear after webhook processing, not immediately after upload. See references/flags.md for ingest options.

Step 5 — Run the benchmark

bash
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
    --data-dir /path/to/datasets --dataset DATASET --subset SUBSET \
    --skip-download --skip-ingest --name RUN_NAME

--skip-ingest reuses an already-indexed dataset; it cannot be combined with --clear or --only-dataset. See references/flags.md before changing flags that affect run meaning.

Decomposition needs no flag. An LLM turns the sentence into {query, attributes, has_action, ...} and that decides which path runs — the same call the deployed agent makes. The script derives the NIM origin from --endpoint; if it cannot reach it the run falls back rather than stopping, and says which fallback it took.

fusion cannot be a fallback. It needs an attribute per query and there is nothing to supply one without decomposition — and feeding it the whole query as its attribute is the failure this eval already measured: mAP −25%, HIT@1 halved, latency +59%.

Leave --concurrency alone unless asked. It defaults to 1 because concurrent queries contend for the same VLM and embedding services, which inflates every per-stage latency in the report.

--name makes the result file findable; without it the name is <dataset>_<subset>_<timestamp>.

Step 6 — Report

Results land in scripts/cli_eval_result/<name>.json.

Report Recall and HIT@k as the quality signal. Precision and mAP are dominated by --top-k — five results against roughly one relevant segment keeps them low however good retrieval is — so they compare runs; they do not grade a deployment.

State these alongside any number, because each one changes what it means:

  • Which paths ran. Search paths: in the summary. A run that was all embed says nothing about attribute.
  • Whether decomposition was live. If so, give its share of query time and what it Routed to: — a routing shift changes what was measured.
  • Whether the critic ran. Critic-filtered metrics print NA when the response carried no verification block. NA is not 0.

Read references/reading-results.md when interpreting stage latencies or metrics. Read references/flags.md when a flag might change comparability. Read references/troubleshooting.md when ingest fails or the CLI exits nonzero. Read references/dataset-format.md when validating a dataset's layout.

Error handling

A non-zero CLI exit aborts the run. A missing result envelope marks that query unanswered and aborts by default; --tolerate-unanswered permits an explicit partial run that excludes those queries from scoring. Never count them as misses. Diagnose an all-zero run with Step 3 before reporting quality. Critic-filtered NA means verification was not available, not zero quality. Use references/troubleshooting.md for the specific failure and recovery steps.

Conventions

  • Never report 0.0000 as a quality result without checking Step 3 first. An unindexed deployment and a broken retriever look identical in the metrics and are not the same finding.
  • A non-zero CLI exit aborts the run rather than scoring 0.0. A stale virtualenv exits 1 on every invocation, and scoring that as "no results" would report a broken environment as an accuracy regression.
  • Queries do not merge adjacent windows by default. Upstream merging averages the scores of merged windows and changes the precision denominator, so a merged run is not comparable to an unmerged baseline.
  • The critic is an LLM and is not deterministic. Raw metrics repeat exactly; critic-filtered ones move a few points between runs. Gate on raw.
  • Live decomposition is not deterministic either. The same query can route differently between runs, so compare Routed to: before comparing scores.

© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 20 other files (scripts, references) in skills/benchmarking/benchmark-video-search of NVIDIA-AI-Blueprints/video-search-and-summarization.

  • SKILL.md
  • .gitignore
  • evals/evals.json
  • references/dataset-format.md
  • references/flags.md
  • references/reading-results.md
  • references/troubleshooting.md
  • scripts/flows/__init__.py
  • scripts/flows/base.py
  • scripts/flows/dataset.py
  • scripts/flows/decompose.py
  • scripts/flows/ingest.py
  • scripts/flows/metrics.py
  • scripts/flows/normalize.py
  • scripts/flows/query.py
  • scripts/flows/readiness.py
  • scripts/flows/routing.py
  • … and 4 more

Open the folder on GitHubat commit fdb6a7a

Compare with similar skills

Benchmark Video Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Video Search compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Video Search this skillNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~4.3kAutomated safety check: PassApache-2.0
Agent BuildershareAI-lab/learn-claude-code78k4 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.9k14 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 4 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 14 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA-AI-Blueprints/video-search-and-summarization

All 22 skills in this repo
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…

    1.9k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Rtvi Vlm Perf Testing

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.

    1.9k GitHub stars~8.6k tokensUpdated yesterday
    Auto-check: notes
  • Vss Build Vision AI

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…

    1.9k GitHub stars~15k tokensUpdated yesterday
    Auto-check: notes
  • Vss Evaluate Caption Accuracy

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

    1.9k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check: notes
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Vss Manage Alerts

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when operating VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, verification verdicts, on-demand verification, always-on operation, Slack…

    1.9k GitHub stars~11k tokensUpdated yesterday
    Auto-check: warnings

Questions about Benchmark Video Search

What does Benchmark Video Search do?

Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths. Benchmark Video Search is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

When should I use Benchmark Video Search?

Benchmark Video Search fits situations like: the user asks to benchmark video search; compare search recall between builds; profile search latency on a deployed VSS profile.

How do I install Benchmark Video Search in Claude Code?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a claude-code`. Or copy the skill folder (skills/benchmarking/benchmark-video-search in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/benchmark-video-search in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Video Search in Codex?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a codex`. Or copy the skill folder (skills/benchmarking/benchmark-video-search in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/benchmark-video-search in your project. Codex loads it when a task matches its description.

Can I use Benchmark Video Search in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-video-search, .gemini/skills/benchmark-video-search, .github/skills/benchmark-video-search and .opencode/skills/benchmark-video-search in your project.

What does Benchmark Video Search need to run?

Going by SKILL.md and its folder, Benchmark Video Search needs Python for the scripts in its folder and the command-line tools its instructions call (uv, curl and python3). Our summary lists: Python 3; Docker.

Does Benchmark Video Search access the network?

SKILL.md contains no URLs. Its commands use uv and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Benchmark Video Search safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Benchmark Video Search use?

Benchmark Video Search is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Video Search use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.9k tokens, read only when the agent opens those files.

What are the alternatives to Benchmark Video Search?

Skills that share tags, products or a category with Benchmark Video Search: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Video Search?

NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.

Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.