Agent skill

Vss Deploy Detection Tracking 3D

by NVIDIA-AI-Blueprints in NVIDIA-AI-Blueprints/video-search-and-summarization

A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC…

Apache-2.0Auto-check: notesAI & LLM Engineering

Install Vss Deploy Detection Tracking 3D

skills CLI
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-deploy-detection-tracking-3d -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-deploy-detection-tracking-3d --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/deployment/vss-deploy-detection-tracking-3d .claude/skills/vss-deploy-detection-tracking-3d && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vss-deploy-detection-tracking-3d
GitHub stars
1.9k
Token cost
~5.1k tokens
SKILL.md length
2,473 words
Files
14 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC…

  • Works in 12 steps: Resolve RTCV3D_APP to… → Identify the input mode: file for local… → If the user asked for the sample dataset… → …
  • The 4-camera sample dataset
  • SKILL.md covers When to Use This Skill, Examples, Output Permissions and What This Deploys, plus 6 more sections
  • Calls docker

What it does

Vss Deploy Detection Tracking 3D is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Use when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC skills, the 4-camera sample dataset, camera config, BEV Fusion, live OSD or saved grid/BEV outputs, bundled brokers, basic external MQTT/Kafka brokers, verification, and teardown. Trigger for generic MV3DT, RTVI-CV-3D, multi-view 3D tracking, multi-cam tracking, or sample MV3DT dataset requests. Explicit warehouse blueprint/profile MV3DT…

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 15 other files, including reference files (for example `BENCHMARK.md`, `evals/calibration-chain.json` and `evals/evals.json`).

It sits in AI & LLM Engineering, covering Deployment, Summarization and Performance reviews. It works with NVIDIA AI Platform and Apache Kafka. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.

When your agent uses it

  • The 4-camera sample dataset
  • Saved grid/BEV outputs
  • Bundled brokers
  • Basic external MQTT/Kafka brokers

Example prompts

  • “/vss-deploy-detection-tracking-3d”

Requirements

  • Docker

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Resolve RTCV3D_APP to services/rtvi/rt-cv-3d/rt-cv-mv3dt.
  2. Identify the input mode: file for local MP4s or stream for RTSP.
  3. If the user asked for the sample dataset or 4-cam example dataset, load references/sample-dataset.md first. Resolve/download app-data, set…
  4. Validate or obtain calibration.json. If missing, hand off to vss-generate-video-calibration by name and do not duplicate the AMC workflow…
  5. Set required values in docker/.env: MODELS_DIR, NUM_CAMS, INPUT_MODE, VIDEO_DIR for file input, and optional image/GPU values. For…
  6. Initialize broker mode before config generation or staging. For bundled mode, run the bundled resource preflight in…
  7. Generate generated/camInfo/ and generated/pub_sub_info_config.yml from calibration.json with the standalone scripts/generate-configs.sh…
  8. Run the concrete display probe from references/configure-cameras.md before staging configs; it must test the current DISPLAY and…
  9. Stage DeepStream configs with scripts/stage-configs.sh, then assert generated/configs/ds-main-config-mv3dt.txt contains a Kafka…
  10. Preflight output/tooling and model-cache writeability. Cold TensorRT engine builds can take 5-10 minutes and must be able to persist…
  11. For every INPUT_MODE=file run, start the selected brokers and bev-fusion, wait for broker/topic-init/BEV Fusion readiness, then capture…
  12. If saved BEV is selected/defaulted, or if file input needs any live/saved BEV visualization, use the two-phase launch in…

What it can do on your machine

Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vss Deploy Detection Tracking 3D loads about 5.1k tokens when it runs, and up to ~40k if it reads all its reference files. Until then it costs about 202 tokens; SKILL.md has 2,473 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~202
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~40k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:76
    pStream REST ports in standalone `docker/.env`. Do not use full-stack `docker compose up -d` as the generic file-mode la
  • NoteMentions a .env fileSKILL.md:101
    ime setup, prerequisites, model/assets, `.env` | `references/deploy-rtvi-cv-3d-stack.md` |
  • NoteMentions a .env fileSKILL.md:120
    5. Set required values in `docker/.env`: `MODELS_DIR`, `NUM_CAMS`, `INPUT_MODE`, `VIDEO_DIR` for file input, and optiona
  • NoteMentions a .env fileSKILL.md:121
    `KAFKA_BOOTSTRAP` are already in `docker/.env`. For external mode, validate broker endpoints and required topics before
  • NoteRuns commands with sudoSKILL.md:129
    e affected model directory, for example `sudo setfacl -m u:<uid>:rwx -m d:u:<uid>:rwx <model-dir>`; do not use broad `ch

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 2,473 words, ~5,089 tokens.

Download SKILL.mdSave it as .claude/skills/vss-deploy-detection-tracking-3d/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
vss-deploy-detection-tracking-3d
description
Use when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC skills, the 4-camera sample dataset, camera config, BEV Fusion, live OSD or saved grid/BEV outputs, bundled brokers, basic external MQTT/Kafka brokers, verification, and teardown. Trigger for generic MV3DT, RTVI-CV-3D, multi-view 3D tracking, multi-cam tracking, or sample MV3DT dataset requests. Explicit warehouse blueprint/profile MV3DT requests route to vss-build-vision-ai; single-camera 2D tracking routes to the 2D tracking or DeepStream skills. Not for full warehouse blueprint deployment, single-camera 2D tracking, camera calibration itself, or VSS summarization, Q&A, and RAG workflows.
license
Apache-2.0
metadata.author
NVIDIA
metadata.version
3.3.0-rc0
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
metadata.tags
nvidia vss rtvi-cv-3d mv3dt multi-camera tracking bev-fusion standalone

VSS Deploy Detection And Tracking 3D

When to Use This Skill

Deploy the standalone RT-CV-3D MV3DT stack from services/rtvi/rt-cv-3d/rt-cv-mv3dt. This is the default path for MV3DT / RTVI-CV-3D / multi-camera tracking requests.

Do not derive MV3DT services from the warehouse blueprint for this skill. Use vss-build-vision-ai only when the user explicitly asks for warehouse MV3DT, the warehouse blueprint, a bp_wh* profile, warehouse compose files, or the combined warehouse application stack. When routing an explicit warehouse MV3DT request, also state the boundary: generic MV3DT uses standalone RT-CV-3D, while warehouse MV3DT uses warehouse/profile deployment. For single-camera 2D detection or tracking, use the 2D tracking or DeepStream skills instead.

Public docs: https://docs.nvidia.com/vss/latest/object-detection-tracking.html.

Examples

Example operation prompts:

  • "Deploy MV3DT on my calibrated four-camera MP4 dataset and save output."
  • "Deploy MV3DT on the sample dataset."
  • "Enable multi-camera tracking on the 4-cam example dataset."
  • "Run RTVI-CV-3D on these RTSP streams after calibration."
  • "Deploy multi-cam tracking; if there is no display, save the videos."
  • "Use an external MQTT broker and external Kafka for this RT-CV-3D deployment."
  • "Verify the standalone RT-CV-3D deployment and show output paths."
  • "Tear down everything for standalone MV3DT."

Output Permissions

Keep output permissions scoped to the standalone runtime paths. If output writes fail, report the directory owner/mode, container user, and relevant logs instead of loosening permissions broadly.

What This Deploys

The standalone compose file is services/rtvi/rt-cv-3d/rt-cv-mv3dt/docker/compose.yml. It deploys:

ServiceContainerRole
perceptionvss-rtvi-cv-mv3dtRT-DETR plus MV3DT DeepStream perception; publishes per-camera 3D measurements to Kafka topic mdx-raw and uses MQTT /trck/* tracklet exchange.
bev-fusionvss-rtvi-cv-bev-fusionConsumes mdx-raw, fuses same-object measurements across cameras, and publishes mdx-bev.
mosquittovss-mosquitto-mv3dtOptional bundled MQTT broker, enabled by the mosquitto compose profile.
kafkakafkaOptional bundled Kafka broker, enabled by the kafka compose profile.
kafka-topic-initkafka-topic-initOptional one-shot topic initializer for mdx-raw and mdx-bev, enabled by the kafka compose profile.

The standalone stack does not deploy VST, VIOS, NvStreamer, Elasticsearch, Kibana, Logstash, video-analytics-api, behavior analytics, SDR controller, warehouse configurator, agents, LLM, or VLM services.

Core Rules

  • Default to bundled brokers and use COMPOSE_PROFILES=mosquitto,kafka for bundled-broker Compose operations. Before generate-configs.sh, stage-configs.sh, or bundled launch, run the bundled resource preflight: reuse existing standalone containers from this app without rewriting ports, reject foreign fixed-name container collisions, and for a fresh start select free Kafka, MQTT, and DeepStream REST ports in standalone docker/.env. Do not use full-stack docker compose up -d as the generic file-mode launch path; file mode must start support services, capture Kafka baselines, optionally prestart BEV, and only then start perception with --no-deps.
  • For explicit external broker requests, collect, export, and validate MQTT_HOST, MQTT_PORT, and KAFKA_BOOTSTRAP; set USE_EXTERNAL_BROKERS=1; generate pub/sub config with MQTT_BROKERS="${MQTT_HOST}:${MQTT_PORT}" ./scripts/generate-configs.sh; verify mdx-raw and mdx-bev already exist on external Kafka with bounded kafka-topics --describe; use external-broker Compose mode without bundled profiles; and verify Kafka offsets against the external KAFKA_BOOTSTRAP. File-mode external-broker runs still follow the same two-phase ordering. Delegate only TLS/auth variants to the standalone README custom-broker section.
  • Require calibrated, time-synchronized multi-camera input. MV3DT needs at least two cameras; 30 FPS sources should be synchronized within about one frame duration.
  • For recorded files, use INPUT_MODE=file; each .mp4 name must match a sensor id in calibration.json and the generated camInfo. File input is a finite batch run: tell the user up front that vss-rtvi-cv-mv3dt exits after end-of-stream and remaining support containers are stopped after successful verification unless the user asks to keep them.
  • For the sample dataset / 4-cam example dataset, load references/sample-dataset.md. Use the standalone sample flow: NGC warehouse app-data for models/videos, repo sample calibration.json and Top.png for calibration/BEV map, generated transforms, INPUT_MODE=file, NUM_CAMS=4, bundled brokers, then the normal display-first visualization decision: live OSD plus live fused BEV when a working display is found and the user did not ask to save; saved grid plus saved fused BEV when headless or explicitly requested.
  • When the user provides MP4 paths, preserve them as deployment inputs. Use their directory as VIDEO_DIR when basenames already match sensor ids; otherwise create generated symlinks named <sensor_id>.mp4 only when the mapping is explicit or unambiguous by count/order. Do not mutate source videos.
  • For live RTSP, use INPUT_MODE=stream. Dynamic REST registration is the first path; stream keys must match the calibration sensor ids. Use the direct REST registration block in references/configure-cameras.md so readiness JSON is parsed independent of whitespace. Do not treat STREAM_ADD_SUCCESS or stream-count alone as success. A live RTSP deployment succeeds only when the expected sources become active, every camera has recent non-zero FPS, and mdx-raw/mdx-bev offsets grow.
  • When the user provides RTSP URLs, preserve them as deployment inputs and register them after the stream-mode compose service is running with the direct REST block in references/configure-cameras.md; the block waits on /api/v1/ready for ds-ready to become YES, so the ds-ready: YES log line is optional diagnostic evidence. Do not stop at telling the user to run registration manually. Ask for mapping only if bare URLs cannot be matched to calibration sensor ids by count/order. If dynamic registration accepts streams but active sources, FPS, or Kafka growth remain zero after bounded verification, treat dynamic add as failed and use the generic static RTSP [source-list] fallback in references/configure-cameras.md with the same user-provided sensor_id=rtsp://... mappings. Do not substitute sample calibration or sample camera mappings unless the user explicitly requested the sample dataset.
  • If calibration is missing, hand off to vss-generate-video-calibration and run its AMC platform preflight before VIOS, capture, upload, or calibration work. If the preflight fails, stop and ask the user to provide existing/generated calibration artifacts or choose a supported x86_64 dGPU/NVENC calibration host. For RTSP calibration, use vss-manage-video-io-storage only to bring up or verify the VIOS prerequisite when VIOS is not already deployed/reachable; AMC owns calibration and VIOS_BASE_URL env wiring once VIOS is available.
  • Do not use VST for visualization. Use the standalone OSD/save-video path and BEV visualizer scripts.
  • Always run the real display probe in references/configure-cameras.md before choosing OSD=0 as the headless fallback; do not infer headless mode only from GPU presence, xdpyinfo installation, or a stale/missing DISPLAY. Treat display mode as two live windows by default: the DeepStream camera-grid OSD and the separate fused BEV visualizer. Treat save video, save output, and confirmed headless fallback as saved perception grid plus saved fused BEV by default. Before launch, preflight host tools needed for selected output: ffprobe for saved artifact verification, and the BEV visualizer Python/OpenCV/Kafka dependencies when BEV visualization/recording is enabled. Before promising BEV, resolve BEV_DATASET_PATH to a directory containing both map.png and transforms.yml; if either is missing, request the missing BEV asset or report perception-grid-only output explicitly.
  • BEV video is not emitted by the perception container. It is produced by the separate host-side scripts/bev-visualizer.sh Kafka consumer. For finite file input, keep the BEV process under the same long-lived shell/session that starts perception, waits for EOS, verifies offsets/artifacts, finalizes BEV, and performs cleanup; do not start BEV in a separate short tool call and assume nohup ... & will survive runner process-group cleanup. Wait for Kafka assignment and verify the PID is still alive immediately before file-mode perception or RTSP stream registration. For finite file-input live display runs, start live fused BEV before perception and, after EOS, tell the user to press q in the BEV window or stop only the tracked current-run BEV PID through the safe teardown flow.

Workflow

Use the workflow selection table and run stages below. Load only the references needed for the user's selected input, broker, visualization, calibration, and verification path.

Workflow Selection

Load the minimum references needed for the current request:

User intentReferences
First-time setup, prerequisites, model/assets, .envreferences/deploy-rtvi-cv-3d-stack.md
Sample dataset, 4-cam example dataset, warehouse 4-camera synthetic datasetreferences/sample-dataset.md, then references/configure-cameras.md, references/deploy-rtvi-cv-3d-stack.md, and references/verify-and-view.md
Existing or newly generated calibration; local MP4 or RTSP input configreferences/configure-cameras.md
Missing calibrationreferences/calibration-workflow.md, then references/configure-cameras.md
Launch or redeploy the stackreferences/deploy-rtvi-cv-3d-stack.md
Add/list/remove live RTSP streamsreferences/configure-cameras.md
Verify containers, logs, Kafka topics, or output artifactsreferences/verify-and-view.md
Live OSD, saved perception video, live BEV, or saved BEV videoreferences/verify-and-view.md
Completed file-input post-run support-service cleanup; stop, tear down everything, or clean generated statereferences/teardown.md
Diagnose failuresreferences/troubleshooting.md
Show full SKILL.md (1,155 more words)Show less

Run Stages

Follow these stages for deployment work:

  1. Resolve RTCV3D_APP to services/rtvi/rt-cv-3d/rt-cv-mv3dt.
  2. Identify the input mode: file for local MP4s or stream for RTSP.
  3. If the user asked for the sample dataset or 4-cam example dataset, load references/sample-dataset.md first. Resolve/download app-data, set MODELS_DIR, VIDEO_DIR=<APP_DATA_DIR>/videos/warehouse-4cams-20mx20m-synthetic, CALIBRATION_JSON, BEV_DATASET_PATH, NUM_CAMS=4, and INPUT_MODE=file, then continue with camera validation and the normal display/save decision before setting OSD, SAVE_VIDEO, or BEV_SAVE_VIDEO.
  4. Validate or obtain calibration.json. If missing, hand off to vss-generate-video-calibration by name and do not duplicate the AMC workflow inline. Explicitly include the AMC platform preflight failure path: stop and request existing/generated calibration artifacts or a supported x86_64 dGPU/NVENC calibration host. After AMC completes, fetch the AMC MV3DT export ZIP, export calibration.json, validate JSON by filtering sensors where type == "camera" and requiring at least two non-empty safe unique camera IDs, then stage BEV assets before continuing. For saved output or BEV viewing, resolve BEV_DATASET_PATH to a directory containing both map.png and transforms.yml before launch.
  5. Set required values in docker/.env: MODELS_DIR, NUM_CAMS, INPUT_MODE, VIDEO_DIR for file input, and optional image/GPU values. For supplied MP4 paths, point VIDEO_DIR at the matching source directory or at a generated symlink directory with one <sensor_id>.mp4 per camera.
  6. Initialize broker mode before config generation or staging. For bundled mode, run the bundled resource preflight in references/deploy-rtvi-cv-3d-stack.md now so selected MQTT_PORT, KAFKA_PORT, and KAFKA_BOOTSTRAP are already in docker/.env. For external mode, validate broker endpoints and required topics before launch.
  7. Generate generated/camInfo/ and generated/pub_sub_info_config.yml from calibration.json with the standalone scripts/generate-configs.sh, using the selected MQTT endpoint; do not mount warehouse MV3DT calibration directories.
  8. Run the concrete display probe from references/configure-cameras.md before staging configs; it must test the current DISPLAY and discovered X socket candidates such as :0/:1, export a working DISPLAY when found, and print RTCV3D_DISPLAY_AVAILABLE. Then choose visualization:
    • If a working display is detected and the user did not ask to save, stage with OSD=1 SAVE_VIDEO=0, set BEV_SAVE_VIDEO=0 BEV_SOURCE=fused, and use live fused BEV visualization when BEV assets are present.
    • If no display is detected, use saved output as the default fallback: set SAVE_VIDEO=1 and save fused BEV after BEV_DATASET_PATH resolves with both required files.
    • If the user asked to save output, set SAVE_VIDEO=1 even when a display exists and also save fused BEV by default after BEV_DATASET_PATH resolves with both required files.
    • If the user asked for both live view and saved output, use OSD=1 SAVE_VIDEO=1 and start saved fused BEV in parallel.
  9. Stage DeepStream configs with scripts/stage-configs.sh, then assert generated/configs/ds-main-config-mv3dt.txt contains a Kafka msg-broker-conn-str matching the selected KAFKA_BOOTSTRAP and RAW_TOPIC. For INPUT_MODE=file, also assert the staged config disables live latency dropping: [source-list] low-latency-mode=0, [source-attr-all] drop-on-latency=0, and [source-attr-all] latency=100000.
  10. Preflight output/tooling and model-cache writeability. Cold TensorRT engine builds can take 5-10 minutes and must be able to persist engines under the mounted model directories. State this warning and the scoped ACL remediation pattern in the user-visible status/report. If a model cache is not writable by the container UID/GID, stop and request approval before applying an ACL to only the affected model directory, for example sudo setfacl -m u:<uid>:rwx -m d:u:<uid>:rwx <model-dir>; do not use broad chmod 777 or broad recursive chown.
  11. For every INPUT_MODE=file run, start the selected brokers and bev-fusion, wait for broker/topic-init/BEV Fusion readiness, then capture Kafka baselines before starting perception. Do this even when saved output or BEV visualization is not requested.
  12. If saved BEV is selected/defaulted, or if file input needs any live/saved BEV visualization, use the two-phase launch in references/deploy-rtvi-cv-3d-stack.md: after support readiness and file baselines, start the BEV visualizer/recorder in the same long-lived shell/session that will start perception, wait for EOS, verify outputs, and finalize BEV. Wait for its Kafka consumer group assignment, verify the recorder PID is still alive, then start perception with --no-deps. Saved output uses BEV_SAVE_VIDEO=1 BEV_SOURCE=fused by default; display-only output uses BEV_SAVE_VIDEO=0 BEV_SOURCE=fused so the BEV window is live. For stream mode with no BEV prestart requirement, full-stack Compose launch is acceptable: bundled uses COMPOSE_PROFILES=mosquitto,kafka docker compose up -d; external uses docker compose up -d. Never use full-stack docker compose up -d for file input.
  13. For RTSP input, after stream-mode compose is running, register the provided streams with the direct REST registration block in references/configure-cameras.md, which waits on /api/v1/ready for ds-ready=YES using JSON parsing. Use explicit <sensor_id>=<rtsp_url> pairs when provided; otherwise map bare URLs to calibration sensor ids only when the counts and ordering are clear. Preserve the final mapping so it can also be used for the static RTSP source-list fallback if dynamic REST add does not produce active sources.
  14. For RTSP, verify REST readiness, exact stream registration, non-zero FPS, mdx-raw/mdx-bev offset growth, and requested visualization artifacts. If registration succeeds but active sources/FPS/Kafka remain zero, restage the same RTSP mapping as a static [source-list], restart only perception as appropriate for stream mode, and rerun the same verification. For file input, do not require ds-ready: YES; treat vss-rtvi-cv-mv3dt Exited (0) with App run successful as EOS success, then require mdx-raw and mdx-bev offsets to be greater than pre-run baselines.
  15. For completed file-input runs, after outputs are verified, stop only the remaining standalone support services unless the user asked to keep them running for inspection or reuse. Preserve generated configs, calibration, videos, and outputs.

Success Criteria

  • generated/camInfo/ contains one .yml per filtered camera sensor and generated/configs/ exists.
  • Runtime images are reported from docker compose config --images; the skill does not infer image tags from its own version.
  • docker compose uses services/rtvi/rt-cv-3d/rt-cv-mv3dt/docker/compose.yml.
  • vss-rtvi-cv-bev-fusion becomes healthy.
  • For RTSP: REST /api/v1/ready reports ds-ready=YES, registered stream count equals NUM_CAMS, registered IDs exactly match generated camInfo IDs with no duplicates/extras, every expected source has recent non-zero FPS, and both mdx-raw and mdx-bev offsets grow while streams are active.
  • For file input: vss-rtvi-cv-mv3dt may end as Exited (0) after EOS and is successful only when logs include App run successful and both mdx-raw and mdx-bev offsets exceed pre-run baselines.
  • If live OSD was selected, display access was checked before staging with OSD=1 without broad xhost +, and display-mode visualization includes both the DeepStream camera-grid OSD window and the separate live fused BEV window when BEV assets are present.
  • If saved output was selected/defaulted, report current-run video-output/grid-view.mkv and saved BEV artifact paths with non-empty size, run-start timestamp checks, ffprobe success, and current BEV log evidence including Video saved with positive frame count. If BEV was skipped because assets were missing, report that explicitly.
  • If live BEV visualization was selected, report the tracked visualizer PID/log, Kafka consumer group assignment evidence, and for finite file input either that the user closed the BEV window with q or that the tracked current-run PID was safely stopped after EOS.
  • vss-generate-video-calibration owns AMC deployment and calibration from local MP4s or RTSP streams.
  • vss-manage-video-io-storage is used only to bring up or verify VIOS when RTSP calibration needs VIOS and it is not already deployed.
  • vss-build-vision-ai owns full warehouse blueprint deployments, including explicit warehouse MV3DT requests.

© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (references) in skills/deployment/vss-deploy-detection-tracking-3d of NVIDIA-AI-Blueprints/video-search-and-summarization.

  • SKILL.md
  • BENCHMARK.md
  • evals/calibration-chain.json
  • evals/evals.json
  • evals/routing.json
  • evals/sample-deployment.json
  • references/calibration-workflow.md
  • references/configure-cameras.md
  • references/deploy-rtvi-cv-3d-stack.md
  • references/sample-dataset.md
  • references/teardown.md
  • references/troubleshooting.md
  • references/verify-and-view.md
  • skill-card.md

Open the folder on GitHubat commit fdb6a7a

Compare with similar skills

Vss Deploy Detection Tracking 3D next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vss Deploy Detection Tracking 3D compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vss Deploy Detection Tracking 3D this skillNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~5.1kAutomated safety check: NotesApache-2.0
Vss Deploy Dense CaptioningNVIDIA/skills3.6k1 repos~3.3kAutomated safety check: NotesApache-2.0
Deepstream Run Mv3dtNVIDIA/skills3.6k—~3.1kAutomated safety check: PassApache-2.0
Monstermq Broker Configvogler75/monster-mq143—~2.2kAutomated safety check: PassGPL-3.0
Deepstream DevNVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.0
Deepstream SopNVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.0

Similar skills

  • Official

    A skill your agent uses when deploying standalone RT-VLM dense captioning or calling its REST API (uploads, captions, streams, chat-completions, Kafka).

    3.6k GitHub starsUsed in 1 repo~3.3k tokens
    Backend & APIsAuto-check: notes
  • Official

    Run and operate the DeepStream Multi-View 3D Tracking reference app, also known as MV3DT.

    3.6k GitHub stars~3.1k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Monstermq Broker Config

    vogler75/monster-mq

    Guide for configuring, deploying, and operating the MonsterMQ broker.

    143 GitHub stars~2.2k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Deepstream Dev

    NVIDIA/skills

    Official

    NVIDIA DeepStream SDK development with Python pyservicemaker API.

    3.6k GitHub stars~3.3k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Deepstream Sop

    NVIDIA/skills

    Official

    A skill your agent uses when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether…

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Backend & APIsAuto-check: notes
  • RAG Blueprint

    NVIDIA/skills

    Official

    NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage.

    3.6k GitHub stars~2.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes

More from NVIDIA-AI-Blueprints/video-search-and-summarization

All 22 skills in this repo
  • Benchmark Video Search

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

    1.9k GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…

    1.9k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Rtvi Vlm Perf Testing

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.

    1.9k GitHub stars~8.6k tokensUpdated today
    Auto-check: notes
  • Vss Build Vision AI

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…

    1.9k GitHub stars~15k tokensUpdated today
    Auto-check: notes
  • Vss Evaluate Caption Accuracy

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

    1.9k GitHub stars~2.1k tokensUpdated today
    Auto-check: notes
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated today
    Auto-check passed

Questions about Vss Deploy Detection Tracking 3D

What does Vss Deploy Detection Tracking 3D do?

A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC…. Vss Deploy Detection Tracking 3D is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Use when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC skills, the 4-camera sample dataset, camera config, BEV Fusion, live OSD or saved grid/BEV outputs, bundled brokers, basic external MQTT/Kafka brokers, verification, and teardown.

When should I use Vss Deploy Detection Tracking 3D?

Vss Deploy Detection Tracking 3D fits situations like: the 4-camera sample dataset; saved grid/BEV outputs; bundled brokers; basic external MQTT/Kafka brokers.

How do I install Vss Deploy Detection Tracking 3D in Claude Code?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-deploy-detection-tracking-3d -a claude-code`. Or copy the skill folder (skills/deployment/vss-deploy-detection-tracking-3d in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/vss-deploy-detection-tracking-3d in your project. Claude Code loads it when a task matches its description.

How do I install Vss Deploy Detection Tracking 3D in Codex?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-deploy-detection-tracking-3d -a codex`. Or copy the skill folder (skills/deployment/vss-deploy-detection-tracking-3d in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/vss-deploy-detection-tracking-3d in your project. Codex loads it when a task matches its description.

Can I use Vss Deploy Detection Tracking 3D in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-deploy-detection-tracking-3d -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vss-deploy-detection-tracking-3d, .gemini/skills/vss-deploy-detection-tracking-3d, .github/skills/vss-deploy-detection-tracking-3d and .opencode/skills/vss-deploy-detection-tracking-3d in your project.

What does Vss Deploy Detection Tracking 3D need to run?

Going by SKILL.md and its folder, Vss Deploy Detection Tracking 3D needs the command-line tools its instructions call (docker). Our summary lists: Docker.

Does Vss Deploy Detection Tracking 3D access the network?

SKILL.md names 1 domain. As links in the text: docs.nvidia.com. This is read from the text; nothing was executed.

Is Vss Deploy Detection Tracking 3D safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Vss Deploy Detection Tracking 3D use?

Vss Deploy Detection Tracking 3D is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vss Deploy Detection Tracking 3D use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 35k tokens, read only when the agent opens those files.

What are the alternatives to Vss Deploy Detection Tracking 3D?

Skills that share tags, products or a category with Vss Deploy Detection Tracking 3D: Vss Deploy Dense Captioning (NVIDIA/skills, 3.6k stars), Deepstream Run Mv3dt (NVIDIA/skills, 3.6k stars), Monstermq Broker Config (vogler75/monster-mq, 143 stars) and Deepstream Dev (NVIDIA/skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vss Deploy Detection Tracking 3D?

NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.

Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.