Official agent skill

Vss Ask Video

by NVIDIA in NVIDIA/skills

A skill your agent uses to ask the VSS agent's videounderstanding tool a fresh visual question about a recorded clip.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Vss Ask Video

skills CLI
$ npx skills add NVIDIA/skills --skill vss-ask-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills vss-ask-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/vss-ask-video .claude/skills/vss-ask-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
vss-ask-video
GitHub stars
3.5k
Token cost
~1.4k tokens
SKILL.md length
545 words
Files
6
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses to ask the VSS agent's videounderstanding tool a fresh visual question about a recorded clip.

  • Works in 3 steps: Probe the VSS agent → If the probe fails, ask the user → If the probe passes, proceed.
  • Ask the VSS agents videounderstanding tool a fresh visual question about a recorded clip
  • SKILL.md covers When to Use, Deployment prerequisite, Sensor prerequisite and Agent workflow, plus 2 more sections
  • Calls curl, jq and python3

What it does

Vss Ask Video is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use this skill to ask the VSS agent's videounderstanding tool a fresh visual question about a recorded clip. Not for prior tool output, search hits, or metadata-answerable questions.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `evals/base_profile_video_understanding.json` and `evals/evals.json`).

It sits in AI & LLM Engineering, covering Computer vision. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Ask the VSS agents videounderstanding tool a fresh visual question about a recorded clip
  • Tasks that involve Computer vision

Example prompts

  • “/vss-ask-video”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Probe the VSS agent
  2. If the probe fails, ask the user
  3. If the probe passes, proceed.

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • jq
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Vss Ask Video loads about 1.4k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 545 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:84
    # Set from deployment (compose / .env / host where vss-agent listens)

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 545 words, ~1,406 tokens.

Download SKILL.mdSave it as .claude/skills/vss-ask-video/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
vss-ask-video
description
Use this skill to ask the VSS agent's video_understanding tool a fresh visual question about a recorded clip. Not for prior tool output, search hits, or metadata-answerable questions.
license
Apache-2.0
metadata.version
3.2.0
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
metadata.tags
nvidia blueprint operational

Video QnA using VLM through VSS Agent

Use this skill when you need details about the video which requires VLM to look at the video frames — for example the agent has no usable prior answer and needs a fresh look at the pixels for a specific clip.


When to Use

  • The user asks what happens in the video, what objects / people / actions appear, colors, timing, safety, or other visual facts that require watching the clip.
  • The user asks for details that cannot be answered from existing messages, summaries, Elasticsearch/MCP results, or filenames alone—you need model inference on the video.
  • Follow-up questions about content details after a coarse summary or after report generation.

Do not use this skill when a database / MCP / prior tool output already answers the question, unless the user explicitly wants verification against the video.


Deployment prerequisite

This skill requires a VSS profile that serves the video_understanding tool — typically base (recommended) or lvs. Before any request:

  1. Probe the VSS agent:

    bash
    curl -sf --max-time 5 "http://${HOST_IP}:8000/docs" >/dev/null
  2. If the probe fails, ask the user:

    "No VSS profile is running on $HOST_IP. Shall I deploy base (recommended for per-clip VLM QnA) using the /vss-deploy-profile skill? If you prefer lvs, say so."

    • If yes → hand off to /vss-deploy-profile -p base (or -p lvs if the user prefers). Return here once it succeeds.
    • If no → stop.
  3. If the probe passes, proceed.


Sensor prerequisite

You MUST list VST sensors before any /generate call. This is required even when the user names the sensor explicitly, even when the user asserts the video is already uploaded, and even when a previous turn appeared to use the same video. Do not skip this step.

  1. List sensors:

    bash
    curl -sf --max-time 5 "http://${HOST_IP}:30888/vst/api/v1/sensor/list" | jq '.[].name'
  2. Compare the returned name values against the user-supplied <sensor-id> (or filename stem, e.g. warehouse_safety_0001).

  3. If a matching sensor is present → proceed to the Agent workflow below.

  4. If no matching sensor is present — upload the video first, then re-list to confirm the new sensor appears:

    bash
    # filename: must not contain whitespace
    # timestamp: ISO 8601 UTC — default 2025-01-01T00:00:00.000Z if user did not specify
    curl -s -X PUT "http://${HOST_IP}:30888/vst/api/v1/storage/file/<filename>?timestamp=<timestamp>" \
      -H "Content-Type: application/octet-stream" \
      -H "Content-Length: <file_size_in_bytes>" \
      --upload-file /path/to/<filename> | jq .

    See /vss-manage-video-io-storage for full upload semantics (v1 vs v2, conflict handling, delete flow). In interactive runs, confirm with the user before uploading. Never issue an unconditional PUT without first running the sensor-list check above — that is exactly the failure mode this prerequisite exists to prevent.


Show full SKILL.md (175 more words)Show less

Agent workflow

The Sensor prerequisite above must have already confirmed (or made) the sensor exist on VST. Then:

  1. Clip — Identify sensor id, filename, or URL for one video segment. If ambiguous, ask the user.
  2. Call vss agent with the sensor id and ask for it to call video_understanding tool to answer the user's question.
  3. Return the vss agent's answer back to the user.

Query VSS agent (/generate)

bash
# Set from deployment (compose / .env / host where vss-agent listens)
export VSS_AGENT_BASE_URL="http://localhost:8000"

curl -s -X POST "${VSS_AGENT_BASE_URL}/generate" \
  -H "Content-Type: application/json" \
  -d '{"input_message": "Call video_understanding tool to answer the following question about <sensor-id>: <user query>"}' | jq .
Response contract and extraction

/generate returns a JSON object with the assistant output in value, for example:

json
{"value":"<agent-think><agent-think-step ...>...</agent-think-step></agent-think>\n\n<final answer>\n\n"}

There is no separate clean-answer field. The consumable answer is the text in .value after removing any <agent-think>...</agent-think> block.

Required handling for this skill (and any downstream caller):

  1. Read .value from the JSON response.
  2. Strip <agent-think>...</agent-think> sections wherever they appear.
  3. Return only the remaining final-answer text to the user.

Example extraction:

bash
curl -s -X POST "${VSS_AGENT_BASE_URL}/generate" \
  -H "Content-Type: application/json" \
  -d '{"input_message":"Call video_understanding tool to answer the following question about <sensor-id>: <user query>"}' \
| jq -r '.value' \
| python3 -c 'import re,sys; t=sys.stdin.read(); t=re.sub(r"<agent-think>.*?</agent-think>\s*", "", t, flags=re.S); print(t.strip())'

Cross-Reference

  • vss-manage-video-io-storage — VST storage/replay URLs so VIDEO_URL is valid for the VLM.
  • vss-generate-video-report — timestamped reports via Mode A (direct VLM) or Mode B (video-analytics incidents); this skill is VSS-agent /generate for ad-hoc video Q&A.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/vss-ask-video of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/base_profile_video_understanding.json
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Vss Ask Video next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Vss Ask Video compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Vss Ask Video this skillNVIDIA/skills3.5k—~1.4kAutomated safety check: NotesApache-2.0
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Yolo Master AgentTencent/YOLO-Master742—~755Automated safety check: PassAGPL-3.0
Video Understandjjyaoao/HelloAgents3.2k1 repos~6.2kAutomated safety check: PassMIT
Motioneyes Visual Analysisedwardsanchez/MotionEyes229—~2kAutomated safety check: PassNone

Similar skills

  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Yolo Master Agent

    Tencent/YOLO-Master

    A skill your agent uses when the user wants to run a YOLO-Master task (train/val/predict/track/export/benchmark) or use the Agent Skill dispatcher.

    742 GitHub stars~755 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Video Understand

    jjyaoao/HelloAgents

    Implement specialized video understanding capabilities using the z-ai-web-dev-sdk.

    3.2k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Motioneyes Visual Analysis

    edwardsanchez/MotionEyes

    Pixel-based motion and UI change analysis from frame sequences or screenshots using computer vision and visual comparison.

    229 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLaVA Vision-Language Model

    Orchestra-Research/AI-Research-SKILLs

    Guide to LLaVA for image chat, visual question answering and captioning, with model sizes, CLI and Gradio usage and multi-turn conversation code.

    13k GitHub starsUsed in 7 repos~2k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Vss Ask Video

What does Vss Ask Video do?

A skill your agent uses to ask the VSS agent's videounderstanding tool a fresh visual question about a recorded clip. Vss Ask Video is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use this skill to ask the VSS agent's videounderstanding tool a fresh visual question about a recorded clip.

When should I use Vss Ask Video?

Vss Ask Video fits situations like: ask the VSS agents videounderstanding tool a fresh visual question about a recorded clip; tasks that involve Computer vision.

How do I install Vss Ask Video in Claude Code?

Run `npx skills add NVIDIA/skills --skill vss-ask-video -a claude-code`. Or copy the skill folder (skills/vss-ask-video in NVIDIA/skills) into .claude/skills/vss-ask-video in your project. Claude Code loads it when a task matches its description.

How do I install Vss Ask Video in Codex?

Run `npx skills add NVIDIA/skills --skill vss-ask-video -a codex`. Or copy the skill folder (skills/vss-ask-video in NVIDIA/skills) into .agents/skills/vss-ask-video in your project. Codex loads it when a task matches its description.

Can I use Vss Ask Video in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill vss-ask-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vss-ask-video, .gemini/skills/vss-ask-video, .github/skills/vss-ask-video and .opencode/skills/vss-ask-video in your project.

What does Vss Ask Video need to run?

Going by SKILL.md and its folder, Vss Ask Video needs the command-line tools its instructions call (curl, jq and python3). Our summary lists: Python 3.

Does Vss Ask Video access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Vss Ask Video safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Vss Ask Video use?

Vss Ask Video is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Vss Ask Video use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Vss Ask Video?

Skills that share tags, products or a category with Vss Ask Video: Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Yolo Master Agent (Tencent/YOLO-Master, 742 stars) and Video Understand (jjyaoao/HelloAgents, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Vss Ask Video?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.