Vss Deploy Profile
NVIDIA/skills
A skill your agent uses to select, configure, deploy, verify, debug, or tear down a VSS profile (base, search, lvs, warehouse, edge).
Agent skill
by NVIDIA-AI-Blueprints in NVIDIA-AI-Blueprints/video-search-and-summarization
Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization rtvi-vlm-perf-testing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/rtvi-vlm-perf-testing .claude/skills/rtvi-vlm-perf-testing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rtvi-vlm-perf-testing" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/rtvi-vlm-perf-testing into .claude/skills/rtvi-vlm-perf-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rtvi-vlm-perf-testing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/rtvi-vlm-perf-testingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization rtvi-vlm-perf-testing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/benchmarking/rtvi-vlm-perf-testing .agents/skills/rtvi-vlm-perf-testing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rtvi-vlm-perf-testing" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/rtvi-vlm-perf-testing into .agents/skills/rtvi-vlm-perf-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rtvi-vlm-perf-testing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization rtvi-vlm-perf-testing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/benchmarking/rtvi-vlm-perf-testing .cursor/skills/rtvi-vlm-perf-testing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rtvi-vlm-perf-testing" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/rtvi-vlm-perf-testing into .cursor/skills/rtvi-vlm-perf-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rtvi-vlm-perf-testing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git --path skills/benchmarking/rtvi-vlm-perf-testing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization rtvi-vlm-perf-testing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/benchmarking/rtvi-vlm-perf-testing .gemini/skills/rtvi-vlm-perf-testing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rtvi-vlm-perf-testing" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/rtvi-vlm-perf-testing into .gemini/skills/rtvi-vlm-perf-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rtvi-vlm-perf-testing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization rtvi-vlm-perf-testingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/benchmarking/rtvi-vlm-perf-testing .github/skills/rtvi-vlm-perf-testing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rtvi-vlm-perf-testing" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/rtvi-vlm-perf-testing into .github/skills/rtvi-vlm-perf-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rtvi-vlm-perf-testing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization rtvi-vlm-perf-testing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/benchmarking/rtvi-vlm-perf-testing .opencode/skills/rtvi-vlm-perf-testing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rtvi-vlm-perf-testing" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/rtvi-vlm-perf-testing into .opencode/skills/rtvi-vlm-perf-testing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rtvi-vlm-perf-testing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rtvi-vlm-perf-testingPlan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.
Rtvi Vlm Perf Testing is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks. Use this skill when running fresh-container stream-capacity, semantic-isolation, latency, throughput, or regression experiments against an RTVI microservices checkout.
Its SKILL.md is about 8.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/performance-contract.md` and `scripts/canary_executor.py`).
It sits in Backend & APIs, covering Deployment and Microservices. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3dockerbashgitrgFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NGC_API_KEYARTIFACTORY_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Rtvi Vlm Perf Testing loads about 8.6k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 4,034 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
by default and ignores stale generated `.env.perf` values for those two keys unless they are explicitly exported in theml`, and exported shell values override `.env.perf`; the full list with defaults is in `perf/benchmark/PERF_GUIDE.RTVI_Vcompose -f compose.perf.yaml --env-file .env.perf downcompose -f compose.perf.yaml --env-file .env.perf up -d`.env.perf` carries no EVS keys; `setup_perf_env.sh` does not write them, but acompose -f compose.perf.yaml --env-file .env.perf exec rtvi-server env | rg 'VIA_EVS|PRUNING' || echo "no EVS vars set"compose -f compose.perf.yaml --env-file .env.perf downcompose -f compose.perf.yaml --env-file .env.perf up -dAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 4,034 words, ~8,588 tokens.
.claude/skills/rtvi-vlm-perf-testing/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.This repository ships the benchmark implementation under services/rtvi/rt-vlm/perf/. Set the manifest repo to that service directory and confirm that it contains perf/benchmark/ before using repository-specific benchmark commands. Read the repository contributor instructions before changing scripts, configs, docs, or benchmark behavior. Check git status --short before editing and preserve unrelated dirty worktree changes.
The committed planning and canary helpers require Python 3.9+. Executed canaries additionally require SSH, tmux, Docker with NVIDIA Container Toolkit, FFmpeg/FFprobe, and an NVIDIA GPU on the named remote host. Model images and protected artifacts may require the user's existing NGC or artifact-registry access; never request that credentials be stored in a manifest.
Do not start or stop services when the user only asks to inspect existing reports. Treat setup, teardown, and benchmark runs as side-effecting operations and state that you are about to run them.
When editing shell scripts, validate with bash -n. When changing commands, environment variables, Docker behavior, benchmark scenarios, or report workflows, update the matching guidance in the repository AGENTS.md, service README, and perf/benchmark/README.RTVI_VLM.md or PERF_GUIDE.RTVI_VLM.md.
Before launching a new experiment, freeze its identity, workload, metrics, scenarios, and isolated
paths using references/performance-contract.md. Validate and render the plan without executing it:
python3 skills/benchmarking/rtvi-vlm-perf-testing/scripts/perf_plan.py validate plan.json
python3 skills/benchmarking/rtvi-vlm-perf-testing/scripts/perf_plan.py render plan.jsonThe helper is standard-library only and never starts workloads. A launch, restart, remote edit, or cleanup remains side effecting and requires the applicable authorization. Use a fresh pinned runtime for every new or rerun experiment; never reuse a stale benchmark container.
For a pinned one-, two-, or four-stream remote GPU canary, use the committed executor instead of recreating a remote
runner. Dry-run is the default; --execute authorizes staging, runtime launch, benchmark, evidence
collection, and cleanup on the host named in the manifest:
python3 skills/benchmarking/rtvi-vlm-perf-testing/scripts/canary_executor.py \
launch /absolute/path/to/canary-manifest.json
python3 skills/benchmarking/rtvi-vlm-perf-testing/scripts/canary_executor.py \
launch /absolute/path/to/canary-manifest.json --executeThe launcher stages one immutable manifest and its three committed helpers, starts one durable
tmux job through a short SSH command, then watches terminal status through a separate connection.
It returns as soon as terminal JSON arrives rather than waiting for SSH EOF. A disconnect leaves the remote job running. The job
writes atomic status.json, append-only events.jsonl, command and service logs, runtime inspection,
results, cleanup evidence, and checksums below <output_root>/<run_id>. Read
references/performance-contract.md before creating its manifest. The executor accepts one, two, four, eight, or sixteen
independent_live_stream sources, plus thirty-two only with checksum-pinned object media. It first validates one declared VST RTSP stream, then starts one MediaMTX publisher per unique RTSP identity and requires fresh
per-stream measurements in every repetition. Provision VST with perf/setup_perf_env.sh; the canary verifies but does not own that runtime. Use the ordinary plan/runner workflow for capacity or
larger multi-stream experiments.
For the two-, four-, eight-, sixteen-, or object-backed thirty-two-stream semantic-leakage gate, set top-level "semantic_isolation": true. The executor
uses deterministic solid colors through eight streams; at sixteen it pairs each color with SOLID or a
contrasting BORDER, sends concurrent caption requests bound to distinct stream IDs, and
requires three source-correct responses from every source before the performance canary starts. It preserves the
stream mapping and captions in evidence/semantic-isolation.json, drains those probe streams, and
fails closed on swapped, mixed, missing, or undeleted results.
To qualify real object fixtures through the same RTSP and model-preprocessing path, add semantic_media
with one safe label per stream, each containing an absolute path and exact sha256. The executor loops
each image through its own publisher, normalizes odd dimensions for H.264, and changes the forced-choice
probe to object recognition. Set "qualification_only": true to stop successfully after this gate; reuse
the unchanged checksum-pinned mapping with false for the subsequent performance run.
Thirty-two streams require this mapping; unqualified synthetic 32-source manifests are rejected.
Before a launch, remove already-present containers only when harness labels prove exact ownership:
python3 skills/benchmarking/rtvi-vlm-perf-testing/scripts/container_guard.py \
--run-id <previous-run-id> --name <expected-container-name> --execute --jsonDerive the Compose project mechanically from the run ID. Label auxiliary containers with
com.nvidia.rtvi.harness.run_id=<run-id>. The guard force-removes matching owned containers when
--execute is authorized, rejects an exact-name unlabeled or mismatched container, and leaves all
unrelated containers untouched. After cleanup, use a new run ID and new output/scratch/cache paths.
The guard requires Python 3.9+ and access to the Docker CLI/socket; it never invokes a shell.
Use the setup help as the source of truth for current flags and environment variables:
bash perf/setup_perf_env.sh -hImportant defaults and requirements:
RTVI_IMAGE defaults to ghcr.io/nvidia-ai-blueprints/vss/vss-rt-vlm:develop-latest, or develop-latest-sbsa when setup detects DGX Spark; override it for pinned or custom images and record the resolved digest for benchmark provenance.NGC_API_KEY is required for NVIDIA container registry access.NVIDIA_VISIBLE_DEVICES selects GPUs for the stack and benchmark.VLM_MODEL_PRESET=cr3-nano-reasoner-fp8 selects CR3 Nano Reasoner FP8 and fills VLM_MODEL_TO_USE=cosmos-reason3 plus MODEL_PATH=ngc:nim/nvidia/cosmos3-nano-reasoner:modelopt-fp8-final_format_fix unless those variables are explicitly exported. VLM_MODEL_PRESET=cr3-nano-reasoner-nvfp4 selects the Blackwell-oriented CR3 Nano Reasoner NVFP4 path ngc:nim/nvidia/cosmos3-nano-reasoner:modelopt-nvfp4-full-quantize-final_format_fix with the same model key.VST_LOCAL_PACKAGE defaults to perf/vst_package.tar.gz; setup prefers it over cached or downloaded VST packages.ARTIFACTORY_USER and ARTIFACTORY_TOKEN are required only when the VST package must be downloaded.nvcr.io/rxczgrvsg8nx/vst-dev/vst-streamprocessing:2.1.0-26.04.1, vst-sensor:2.1.0-26.04.1, vst-ingress:2.1.0-26.04.1, and nvstreamer:2.1.0-26.04.1; override with VST_IMAGE_TAG, VST_IMAGE_REGISTRY, or per-image variables.VLLM_ENABLE_PREFIX_CACHING=false and VLLM_DISABLE_MM_PREPROCESSOR_CACHE=true. perf/setup_perf_env.sh writes these values by default and ignores stale generated .env.perf values for those two keys unless they are explicitly exported in the shell for a non-standard experiment.RTVI_ENABLE_GOP_DECODE_OPT defaults to true and only affects file-based decoding. It attaches a GOP-aware probe that skips delta frames for GOPs without selected target timestamps; disable with false, 0, no, or off when isolating file-decode behavior. It has no effect on live RTSP or when all frames are selected.VLLM_MM_ENCODER_ATTN_BACKEND empty for the patched default vLLM path. Use XFORMERS only as an explicit experiment or emergency workaround if a current-vLLM image cannot be rebuilt with the Qwen2.5-VL vision attention patch.VLLM_IGNORE_EOS. Use true only when the intended measurement requires fixed-length generation up to max_tokens; do not compare those runs directly with runs that allow EOS to stop generation early.VIA_EVS_SESSION=true enables EVS++ session mode with similarity-based pruning; VLM_VIDEO_PRUNING_RATE on its own enables fixed-rate pruning. Both are service-level settings read only by the RTVI VLM container, so changing them requires redeploying the stack. They cannot be set per benchmark request or per scenario.VLM_VIDEO_PRUNING_RATE=0.5 whenever EVS++ is enabled. This is what activates pruning inside vLLM: the engine only enables its EVS path when video_pruning_rate is passed at engine init, and this variable is its only source. With VIA_EVS_SESSION=true but the rate unset, sessions are still created and the handler still carries its internal 0.5, but the engine never prunes, so the run pays EVS++ overhead with none of the token reduction. Treat a session-mode run with unchanged prompt-token counts as this misconfiguration until proven otherwise.VIA_EVS_SESSION=true raises ValueError at startup on NemotronH_Nano_Omni_Reasoning_V3 instead of falling back, so a failed service start after enabling EVS on that architecture is expected behavior rather than a deployment fault.VLLM_IGNORE_EOS=true to run to max_tokens. Without it they stop at the natural EOS and report far fewer generated tokens than the non-EVS path, which reads as a throughput win when it is really a shorter OSL.VIA_EVS_SESSION, VLLM_EVS_SIMILARITY_THRESHOLD, and VIA_EVS_TOKEN_BUDGET in every report. The standard EVS++ perf configuration is VLLM_EVS_SIMILARITY_THRESHOLD=0.4 and VIA_EVS_TOKEN_BUDGET=1; flag any run that deviates before comparing it against a baseline. All EVS variables are passed through by compose.yaml and compose.perf.yaml, and exported shell values override .env.perf; the full list with defaults is in perf/benchmark/PERF_GUIDE.RTVI_VLM.md.max_model_len, processor/backend settings, and KV cache dtype. FP8 model weights do not by themselves prove that KV cache is FP8.RTVI_EMPTY_CUDA_CACHE_ON_RESULT=false for perf runs. Enabling per-result torch.cuda.empty_cache() is diagnostic-only because it adds allocator synchronization on the hot path.RTVI_RTSP_LATENCY=300, RTVI_RTPJITTERBUFFER_DROP_ON_LATENCY=false, RTVI_RTPJITTERBUFFER_FASTSTART_MIN_PACKETS=2, RTVI_DISABLE_LIVESTREAM_PREVIEW=true, and RTVI_ENABLE_LIVE_TIMESTAMP_FILTER=false unless the user is intentionally testing those knobs.nvidia-smi dmon; avoid regular nvidia-smi polling in perf loops.Expected benchmark videos live under the configured PERF_VIDEOS_DIR. Setup fetches the LVS warehouse source from NGC and derives missing 1080p, 10 FPS benchmark clips with FFmpeg.
Run setup from the repo root:
bash perf/setup_perf_env.shThe setup script prepares VST, nvstreamer, Redis, Prometheus, Grafana, and RTVI VLM; patches VST RTSP URLs into benchmark configs; and starts the monitoring stack. Use teardown after runs:
bash perf/teardown_perf_env.shIf setup fails, first inspect missing environment variables, Docker login or image pull failures, VST package availability, missing video downloads, and RTSP endpoint substitutions in perf/benchmark/rtvi_vlm_config_*.yaml.
Activate the perf virtual environment when present:
source ~/rtvi-vlm-perf-env/bin/activateUse the platform-specific config that matches the machine under test:
perf/benchmark/rtvi_vlm_config_h100.yamlperf/benchmark/rtvi_vlm_bcd_3_2_config.yamlperf/benchmark/rtvi_vlm_config_l40s.yamlperf/benchmark/rtvi_vlm_config_rtx_pro.yamlperf/benchmark/rtvi_vlm_config_jetson.yamlperf/benchmark/rtvi_vlm_config_spark.yaml or rtvi_vlm_config_test.yamlTypical command shape:
python3 perf/benchmark/rtvi_perf_benchmark.py --config perf/benchmark/rtvi_vlm_config_h100.yaml --scenario max_live_streams_test_1_token_448Before launching a named scenario, especially from copied instructions, validate the exact names in the selected config:
python3 perf/benchmark/rtvi_perf_benchmark.py --config <config.yaml> --list-scenariosRun the scenario families requested by the user or needed for comparison:
max_live_streams_test_1_token, max_live_streams_test_100_token, max_live_streams_test_1_token_448, and max_live_streams_test_100_token_448.single_stream or concurrency.file_burst or e2e_latency.Useful max-stream overrides include --initial-stream-count, --add-stream-count, --binary-search-refinement, and --no-binary-search-refinement. Useful concurrency and latency overrides include --concurrency-levels.
For BCD 3.2, run the named scenarios from rtvi_vlm_bcd_3_2_config.yaml:
max_live_streams_test_1_token_2k, max_live_streams_test_100_token_2k, max_live_streams_test_1_token_4k, max_live_streams_test_100_token_4k, max_live_streams_test_1_token_8k, and max_live_streams_test_100_token_8k.concurrency_test_1_token_2k, concurrency_test_100_token_2k, concurrency_test_1_token_4k, concurrency_test_100_token_4k, concurrency_test_1_token_8k, and concurrency_test_100_token_8k.file_burst_1_token_2k, file_burst_100_token_2k, file_burst_1_token_4k, file_burst_100_token_4k, file_burst_1_token_8k, and file_burst_100_token_8k.e2e_latency_1_token_2k, e2e_latency_100_token_2k, e2e_latency_1_token_4k, e2e_latency_100_token_4k, e2e_latency_1_token_8k, and e2e_latency_100_token_8k.For each configured BCD 3.3 platform, monitor the max-live initial load and stability window. If a frozen latency gate aborts the run before normal refinement, preserve the failed count, logs, drops, stream/source counts, and cleanup proof; then retry with initial_stream_count halved (round down) under a fresh run ID and isolated outputs. Keep the platform, source/image/model, workload shape, graph/IPC mode, thresholds, and seed fixed. Repeat until the initial window passes, then let the runner find the highest-stable/first-unstable boundary. Stop as inconclusive if one stream fails, owned cleanup cannot be proved, or a fatal hardware/runtime gate fires. Track every attempt and the active monitor per platform; a latency breach during normal refinement is the measured unstable point, not a reason to restart. Neither a lower-count retry nor initial launches erase the original BCD target failure.
max-live, concurrent-live, or file-burst), counted load unit, synchronized or staggered start policy, source identity policy, and session reuse policy. Do not compare stream counts across modes as equivalent capacity.For H100 BCD 3.2 full-suite reruns, seed max-live-stream probes near the last
clean H100 reference to reduce convergence time. The reference run
rtvi-vlm-bcd32-full-sessionreset-rtsp-20260504T191256Z found max streams of
160/115 at 2K, 78/58 at 4K, and 36/28 at 8K for OSL=1/100 respectively. Use
these conservative starts unless the hardware, model, RTSP source, or BCD
settings changed:
max_live_streams_test_1_token_2k: initial_stream_count=144, add_stream_count=4max_live_streams_test_100_token_2k: initial_stream_count=104, add_stream_count=4max_live_streams_test_1_token_4k: initial_stream_count=70, add_stream_count=2max_live_streams_test_100_token_4k: initial_stream_count=52, add_stream_count=2max_live_streams_test_1_token_8k: initial_stream_count=32, add_stream_count=1max_live_streams_test_100_token_8k: initial_stream_count=25, add_stream_count=1For BCD file-burst sweeps, record every concurrency level before judging the scenario. A request failure at an overloaded high level, such as 128 concurrency, can make the whole test case report failed even when lower levels are valid. Preserve the last clean concurrency level, its p95/p99 latency, GPU mean, and the exact failure text so the report can distinguish throughput limit from service crash.
For long full-suite BCD runs, create a timestamped temporary config that only changes output_dir, and tee logs under /tmp/rtvi-bcd-runs/. Monitor progress with a filtered summary instead of draining huge per-stream statistics:
rg "Stability check|Adding [0-9]+ stream|System degradation|UNSTABLE|Scenario|Error|FAILED|completed|Phase 2" /tmp/rtvi-bcd-runs/<run>.log | tail -n 40Record the three run handles before walking away from a long benchmark:
CONFIG: timestamped config under /tmp/rtvi-bcd-runs/REPORT: timestamped report directory under the repo rootLOG: tee log under /tmp/rtvi-bcd-runs/When monitoring an active long run, prefer tail or scenario-scoped awk over
polling the PTY session. If the log is quiet, check the log mtime and the
benchmark process before declaring a hang; max-stream windows can be separated
by 60 seconds or more, and freshness waits can intentionally run for several
minutes. If the mtime is stale beyond the configured stability or freshness
timeout, inspect RTVI/VST container logs and active stream count before
interrupting the benchmark.
For BCD e2e_latency_* file cases, the 60-minute videos can legitimately keep
the benchmark log quiet for many minutes, especially with OSL=100. Check the
per-video request_timeout_seconds, RTVI health, recent /generate_captions
or /chat/completions server logs, and server/vLLM worker activity before
calling it a hang. Short nvidia-smi dmon windows can show 0% SM/NVDEC during
CPU-side file decode or request scheduling; prefer the final DCGM summary for
the completed request. For live diagnosis, use a short-window tool such as
top -b -d 2 -n 3 -p <pids>; ps %CPU on long-lived RTVI/vLLM workers is a
lifetime average and can falsely suggest current CPU activity. If the file
upload has no paired generation response, the benchmark socket remains open,
short-window CPU is idle, and dmon is flat beyond the request timeout, treat
the request as a stuck server-side file path and preserve logs before restart.
If a live-stream run stops at Cleaning up N active streams..., inspect RTVI
server logs before calling it a benchmark hang. Current benchmark code should
prefer DELETE /v1/streams/delete-batch for cleanup; when that batch request
reaches the server it should complete in one drain-timeout window instead of
roughly N * drain_timeout. If logs show per-stream
DELETE /v1/streams/delete/<id> calls, or the active benchmark was launched
before the batch-cleanup patch, treat it as old sequential cleanup behavior.
Look for Drain timed out after ...; forcing completion followed by DELETE 200
responses. That indicates teardown is progressing slowly, not that latency
measurement is still running.
If source or config changes are committed while a benchmark is already running, call out that the active process is using the code and generated config from launch time. Let the active run finish if it is producing valid data; use a new setup or benchmark launch for measurements that must include the latest changes.
For RTVI VLM scenarios, use vlm_api_mode: chat_completions to benchmark
/v1/chat/completions instead of /v1/generate_captions. Configure
chat_completions_params for chat-specific generation overrides and
chat_messages for explicit OpenAI-format system/user/assistant messages.
The default remains generate_captions, and existing generate_captions_params
are reused when chat-specific params are omitted.
For EVS++ runs, redeploy to change settings because the benchmark CLI cannot change them.
Run the non-EVS baseline as its own deployment, and clear the EVS variables
explicitly instead of relying on shell state. Compose interpolates
${VIA_EVS_SESSION:-}, so an export left over from an earlier EVS run silently turns
the baseline into an EVS run. Keep VLLM_IGNORE_EOS=true so baseline OSL still matches
the EVS run:
cd docker
unset VIA_EVS_SESSION VLM_VIDEO_PRUNING_RATE VLLM_EVS_SIMILARITY_THRESHOLD VIA_EVS_TOKEN_BUDGET
export VLLM_IGNORE_EOS=true
docker compose -f compose.perf.yaml --env-file .env.perf down
docker compose -f compose.perf.yaml --env-file .env.perf up -dunset only clears the shell. --env-file also feeds interpolation, so confirm
.env.perf carries no EVS keys; setup_perf_env.sh does not write them, but a
hand-edited file can. Verify the baseline container has no EVS settings before
benchmarking it:
docker compose -f compose.perf.yaml --env-file .env.perf exec rtvi-server env | rg 'VIA_EVS|PRUNING' || echo "no EVS vars set"Then deploy the EVS++ run with this configuration:
cd docker
export VIA_EVS_SESSION=true VLM_VIDEO_PRUNING_RATE=0.5 VLLM_IGNORE_EOS=true
export VLLM_EVS_SIMILARITY_THRESHOLD=0.4
export VIA_EVS_TOKEN_BUDGET=1
docker compose -f compose.perf.yaml --env-file .env.perf down
docker compose -f compose.perf.yaml --env-file .env.perf up -dVIA_EVS_TOKEN_BUDGET=1 is the default; it forces generation per clip instead of
accumulating visual tokens across clips, holding caption counts equal to the
baseline so latency and throughput compare directly.
Re-run identical scenario names for the baseline and the EVS++ run so report rows
align. All benchmark modes exercise EVS++ when session mode is on, because routing
depends on VIA_EVS_SESSION alone; the file-based modes go through the session path
just as the live-stream modes do.
When asked to improve throughput or isolate a regression, keep the measurement baseline and optimization experiment separate. Baseline first with cache-disabled benchmark settings, BCD no-drop RTSP settings, a timestamped config, a timestamped report directory, and a tee log.
Look for avoidable host synchronization and CPU/GPU copies before changing model behavior:
.cpu(), .numpy(), .item(), torch.cuda.synchronize(), torch.cuda.empty_cache(), blocking queue waits, per-frame logging, per-chunk subprocess calls, and per-request upload/delete work inside measured steady state.nvidia-smi polling to the measurement loop.nvidia-smi dmon or DCGM samples before calling it under-saturation. Ten-second RTSP chunks can create decode and inference bursts even when the final aggregate GPU mean is high.For CPU/GPU transfer reduction, use this roadmap:
RTVI_EMPTY_CUDA_CACHE_ON_RESULT=false, pre-upload file-burst media, and avoid copying tensors to CPU only for bookkeeping.VLLM_MM_TENSOR_IPC=torch_shm or the legacy VLLM_MULTIMODAL_TENSOR_IPC=true. Verify at vLLM startup that the engine argument is accepted; older vLLM builds silently keep the CPU-copy path when they ignore guarded args.Benchmark outputs usually land in rtvi-vlm-perf-report or a custom dated report directory. Inspect JSON before generating summaries: execution_summary.json, max_live_streams_results.json, test_case_summary.json, and scenario-specific result files.
Generate charts and XLSX from perf/benchmark:
cd perf/benchmark
python3 plot_perf_reports.py all --reports h100=./rtvi-vlm-perf-report --configs h100=rtvi_vlm_config_h100.yaml --output ./perf_charts
python3 generate_perf_xlsx.py --reports H100=./rtvi-vlm-perf-report --configs H100=rtvi_vlm_config_h100.yaml --charts ./perf_charts --output perf_report.xlsx --release "3.1 EA2"Use openpyxl to inspect XLSX files when available. If it is unavailable, read workbook XML from the XLSX zip as a fallback. Do not commit generated report directories, XLSX files, or run-specific hardware outputs unless the user explicitly asks for an artifact to be checked in.
BCD XLSX rows should expose min, max, avg, p50, p75, p90, p95, and p99 latency where the report JSON records them. They should also expose per-stage chunk latency for decode, queue, VLM inference, server processing, and server E2E when request profiling is enabled. Check dropped chunks for max-stream rows and failed request/error-rate columns for file-burst and request-latency rows before calling a BCD run clean.
Before committing after setup_perf_env.sh or a live benchmark, explicitly check for generated artifacts and machine-specific config substitutions:
git status --short
git diff -- perf/benchmark/rtvi_vlm_config_*.yaml perf/benchmark/rtvi_vlm_bcd_3_2_config.yamlDo not commit hard-coded lab RTSP URLs, local backend ports, generated report directories, memory logs, XLSX files, or hardware metric outputs. Keep RTSP_STREAM_URL placeholders in checked-in template configs unless the user explicitly requests a committed machine-specific config.
When comparing reports, align results by platform, scenario, token budget, resolution, model, and stream source. Report max-stream deltas first, then latency percentiles, throughput, GPU memory, power, NVDEC utilization, and failure counts when available.
Flag likely non-code causes before attributing regressions: changed GPU count, driver or container versions, model changes, VST image tag drift, hard-coded RTSP endpoints, missing or different benchmark videos, stale generated report files, and monitor gaps.
For max-stream runs, check that phase2_final_stable and phase2_unstable_ceiling are coherent. If linear probing reaches the cap without instability, an old unstable ceiling can make the report look contradictory. For BCD max live streams, a stream count is only meaningful when fresh stream coverage meets the threshold, dropped chunks are zero, and p90/p95 latency stays below the 10 second chunk real-time limit. Leave enable_latency_growth_instability_check disabled for BCD no-drop measurements; enable it only for a diagnostic run where consecutive latency growth itself should fail the probe. Instantaneous GPU/NVDEC samples are bursty with 10 second RTSP chunks; use aggregate report metrics before concluding that GPU is or is not saturated.
When diagnosing inconsistent live-stream latency:
Fresh streams: 0/N with 0 latency means the probe has no new responses yet, not a low-latency stable state. Inspect RTVI container logs, stream add success, backend URL, and whether the server is producing SSE events.concurrent_live_streams report stats now discard leading same-stream startup burst samples by default. These are queued 10 second RTSP chunks emitted within a short local time window after a delayed stream start; they are recorded in latency_filter and excluded from final avg/p90/p95/p99. Sustained high latencies that arrive at normal chunk cadence are still included and indicate queueing or saturation.concurrency_test_*, each concurrency level runs for a fixed duration. Do not judge from ramp-up samples alone; wait for the final p95/GPU/NVDEC summary. Sustained rising per-stream latency above the 10 second chunk budget at 64 or 128 streams indicates queueing or saturation even when stream creation had zero errors.GPU latest % and NVDEC latest % are single burst samples. Use the final Prometheus/DCGM means for utilization claims, especially with 10 second RTSP chunks.rtpjitterbuffer. If RTVI_RTPJITTERBUFFER_FASTSTART_MIN_PACKETS appears ineffective, inspect container logs or GST debug for the actual element properties.Pipeline disposal timed out and CUDA OOM after repeated probes usually point to teardown, buffer lifetime, or decoder cache handoff problems; inspect video_file_frame_getter.py and container logs.Cleaning up N active streams... often mean sequential stream deletion is waiting on RTVI live-stream drain timeouts. Confirm with RTVI logs before interrupting; a DELETE 200 every drain timeout means cleanup is still advancing./files upload/delete is outside measured steady state.VIA_EVS_TOKEN_BUDGET=1 specifically so counts match; if the EVS run emitted far fewer captions, budget accumulation or generation gating is still active, and the throughput numbers are not comparable rather than improved.VLLM_EVS_SIMILARITY_THRESHOLD prunes more frames, so caption coverage and quality need a sanity check alongside the perf delta.When this skill itself changes, update both the repo copy under
skills/benchmarking/rtvi-vlm-perf-testing/SKILL.md and the installed Codex copy under
~/.codex/skills/rtvi-vlm-perf-testing/SKILL.md when that local copy exists.
After changing the plan or canary contract, run
(cd skills/benchmarking/rtvi-vlm-perf-testing/scripts && python3 -m unittest -v test_perf_plan.py test_canary_executor.py && python3 -m py_compile perf_plan.py container_guard.py canary_executor.py).
Finish with the commands run, report paths inspected or generated, clear regression findings, and any validation that was skipped because it would require starting services or running long benchmarks.
© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in skills/benchmarking/rtvi-vlm-perf-testing of NVIDIA-AI-Blueprints/video-search-and-summarization.
Open the folder on GitHubat commit fdb6a7a
Rtvi Vlm Perf Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Rtvi Vlm Perf Testing this skillNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~8.6k | Automated safety check: Notes | Apache-2.0 | |
| Vss Deploy ProfileNVIDIA/skills | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | |
| Rtvi Vlm Customize ModelNVIDIA/skills | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | |
| Frontmcp Deploymentagentfront/frontmcp | 146 | — | ~9.2k | Automated safety check: Notes | Apache-2.0 | |
| AWS Cloudformation Lambdagiuseppe-trisciuoglio/developer-kit | 357 | — | ~3k | Automated safety check: Notes | MIT | |
| Linkerd Patternswshobson/agents | 40k | 9 repos | ~1.8k | Automated safety check: Pass | MIT |
NVIDIA/skills
A skill your agent uses to select, configure, deploy, verify, debug, or tear down a VSS profile (base, search, lvs, warehouse, edge).
NVIDIA/skills
How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.
agentfront/frontmcp
A skill your agent uses when deploying, building for production, packaging, or shipping a FrontMCP server.
giuseppe-trisciuoglio/developer-kit
Provides AWS CloudFormation patterns for Lambda functions, layers, API Gateway integration, event sources, cold start optimization, monitoring, logging, template validation, and deployment workflows.
wshobson/agents
Implement Linkerd service mesh patterns for lightweight, security-focused service mesh deployments.
sickn33/agentic-awesome-skills
Use service mesh patterns for AI inference traffic management, mTLS, canary releases, policy enforcement, and cross-cluster resilience.
NVIDIA-AI-Blueprints/video-search-and-summarization
Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…
NVIDIA-AI-Blueprints/video-search-and-summarization
Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…
NVIDIA-AI-Blueprints/video-search-and-summarization
Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when operating VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, verification verdicts, on-demand verification, always-on operation, Slack…
Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks. Rtvi Vlm Perf Testing is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.
Rtvi Vlm Perf Testing fits situations like: running fresh-container stream-capacity; semantic-isolation; regression experiments against an RTVI microservices checkout.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a claude-code`. Or copy the skill folder (skills/benchmarking/rtvi-vlm-perf-testing in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/rtvi-vlm-perf-testing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a codex`. Or copy the skill folder (skills/benchmarking/rtvi-vlm-perf-testing in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/rtvi-vlm-perf-testing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-vlm-perf-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rtvi-vlm-perf-testing, .gemini/skills/rtvi-vlm-perf-testing, .github/skills/rtvi-vlm-perf-testing and .opencode/skills/rtvi-vlm-perf-testing in your project.
Going by SKILL.md and its folder, Rtvi Vlm Perf Testing needs Python for the scripts in its folder, the command-line tools its instructions call (python3, docker, bash, git and rg) and credentials named NGC_API_KEY and ARTIFACTORY_TOKEN. Our summary lists: Python 3; Docker; A credential in NGC_API_KEY; A credential in ARTIFACTORY_TOKEN.
SKILL.md contains no URLs. Its commands use docker and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Rtvi Vlm Perf Testing is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 8.6k tokens (SKILL.md is roughly 34k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Rtvi Vlm Perf Testing: Vss Deploy Profile (NVIDIA/skills, 3.6k stars), Rtvi Vlm Customize Model (NVIDIA/skills, 3.6k stars), Frontmcp Deployment (agentfront/frontmcp, 146 stars) and AWS Cloudformation Lambda (giuseppe-trisciuoglio/developer-kit, 357 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.
Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.