Official agent skill

Rtvi Vlm Customize Model

by NVIDIA in NVIDIA/skills

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

OfficialApache-2.0Auto-check: notesBackend & APIs

Install Rtvi Vlm Customize Model

skills CLI
$ npx skills add NVIDIA/skills --skill rtvi-vlm-customize-model -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills rtvi-vlm-customize-model --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rtvi-vlm-customize-model .claude/skills/rtvi-vlm-customize-model && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rtvi-vlm-customize-model
GitHub stars
3.5k
Token cost
~5k tokens
SKILL.md length
1,576 words
Files
9 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

  • Tasks that involve Microservices
  • SKILL.md covers When to use, Instructions, Examples and Source locations, plus 7 more sections
  • Calls docker, curl and jq; reaches huggingface.co and integrate.api.nvidia.com; needs VIA_VLM_API_KEY and RTVI_VLM_API_KEY
  • Tasks that involve Deployment

What it does

Rtvi Vlm Customize Model is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/health-checks.md`).

It sits in Backend & APIs, covering Microservices, Deployment and LLM inference and serving. It works with NVIDIA AI Platform, OpenAI and vLLM. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Microservices
  • Tasks that involve Deployment
  • Tasks that involve LLM inference and serving

Example prompts

  • “/rtvi-vlm-customize-model”

Requirements

  • Docker
  • A credential in RTVI_VLM_API_KEY
  • A credential in VIA_VLM_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • curl
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • huggingface.co
    • integrate.api.nvidia.com
    • api.openai.com

    Also links to:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VIA_VLM_API_KEY
    • RTVI_VLM_API_KEY
    • OPENAI_API_KEY
    • HF_TOKEN
    • NGC_CLI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rtvi Vlm Customize Model loads about 5k tokens when it runs, and up to ~7.8k if it reads all its reference files. Until then it costs about 52 tokens; SKILL.md has 1,576 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~52
When it runs · the whole SKILL.md, loaded when a task matches
~5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:58
    `RTVI_VLM_*` vars in `${VSS_PROFILE_DIR}/.env` / `generated.env` |
  • NoteMentions a .env fileSKILL.md:60
    _URL`, `VLM_NAME` in `${VSS_PROFILE_DIR}/.env` |
  • NoteMentions a .env fileSKILL.md:94
    The VSS blueprint `.env` / `generated.env` maps most `RTVI_VLM_*` variables to container-native names. The model or depl
  • NoteMentions a .env fileSKILL.md:99
    profile `.env`, then rerun `dev-profile.sh`. Edit `.env`, not `generated.env` —
  • NoteMentions a .env fileSKILL.md:103
    # ${VSS_PROFILE_DIR}/.env
  • NoteMentions a .env fileSKILL.md:130
    # ${VSS_PROFILE_DIR}/.env
  • NoteMentions a .env fileSKILL.md:240
    # services/rtvi/rt-vlm/docker/.env
  • NoteMentions a .env fileSKILL.md:342
    # from .env
  • NoteMentions a .env fileSKILL.md:343
    tp://<host-ip>:${RTVI_VLM_PORT}   # from .env
  • NoteMentions a .env fileSKILL.md:348
    efault. Define it in the Alerts profile `.env` and regenerate

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,576 words, ~4,983 tokens.

Download SKILL.mdSave it as .claude/skills/rtvi-vlm-customize-model/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
rtvi-vlm-customize-model
description
How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.
metadata.version
1.0.0
metadata.team
accelerated-microservices
metadata.author
NVIDIA CORPORATION <info@nvidia.com>
metadata.tags
vlm, vss, rtvi
metadata.languages
en
metadata.domain
computer-vision

VLM Customization — VSS Alerts Blueprint

When to use

Use this skill when the user wants to:

  • repoint the VSS Alerts Blueprint to a different VLM endpoint,
  • run rtvi-vlm standalone with either an OpenAI-compatible endpoint or the in-container vLLM path,
  • fix the assumption that changing one VLM config automatically updates rtvi-vlm, vlm-as-verifier, and vss-agent.

Do not use this skill for CV detector swaps inside vss-rt-cv; use rtvi-cv-customize-model for those.

Instructions

  • Treat rtvi-vlm, vlm-as-verifier, and vss-agent as three separate VLM consumers. Do not imply that changing only RTVI_VLM_* automatically repoints the verifier or the agent UI.
  • Treat all paths in this skill as relative to a checkout of the public VSS Blueprint repository. Blueprint deployment files live under deploy/docker/...; customer-accessible RT-VLM source and standalone deployment files live under services/rtvi/rt-vlm/.... None of these paths resolve inside the DeepStream repository.
  • For blueprint-side OpenAI-compatible routing, cover all three surfaces: RTVI_VLM_* env vars, ${VSS_PROFILE_DIR}/vlm-as-verifier/configs/config.yml, and the vss-agent VLM_MODEL_TYPE / VLM_NAME / VLM_BASE_URL settings, then mention force-recreate plus log checks.
  • For standalone vllm-compatible, set VLM_MODEL_TO_USE=vllm-compatible, point MODEL_PATH at the weights source, mention HF_TOKEN only when the model actually needs auth. Do not treat a log line as proof of a working deployment: v3.2.1 logs Warmup VlmProcess-0 done even when warm-up raised an exception. Require the absence of an Error during warmup line, readiness, and one successful /v1/chat/completions response.
  • Never paste credentials into chat or logs; capture with a silent prompt (read -rsp), chmod 600 the env file, and never commit generated.env.

Examples

  • "Point the VSS Alerts Blueprint at a host-side NIM on http://host.docker.internal:30082 for all RTVI-VLM calls."
  • "Run rtvi-vlm standalone with Qwen3-VL-8B-Instruct served inside the container."
  • "I changed only RTVI_VLM_ENDPOINT in generated.env. That should repoint vlm-as-verifier and vss-agent too, right?"

Source locations

This skill is documentation-only — no VSS sources ship here. Use a VSS v3.2.1 or compatible checkout and run every command from its root. Clone/LFS, Alerts profile paths, RT-VLM tree, pinned images, and VSS_* vars: references/vss-source-layout.md.

Which modes this applies to

Mode flagWorkflowVLM role
--mode real-time (2d_vlm)Real-time VLM alertsPrimary — every alert is VLM-driven
--mode verification (2d_cv)CV + VLM verificationVerifier — VLM confirms each CV incident

VLM customization applies to both modes. Three services each make their own VLM calls with separate configuration:

ServiceWhen usedConfig location
rtvi-vlmReal-time alert generation; rtvi_vlm_alert tool calls from agent UIRTVI_VLM_* vars in ${VSS_PROFILE_DIR}/.env / generated.env
vlm-as-verifier (alert-bridge)Post-processing: confirms each mdx-incidents event (2d_cv only)${VSS_PROFILE_DIR}/vlm-as-verifier/configs/config.yml
vss-agentInteractive agent UI queriesVLM_MODEL_TYPE, VLM_BASE_URL, VLM_NAME in ${VSS_PROFILE_DIR}/.env

All three can point at the same model endpoint — they don't have to.

alert-bridge and vss-agent are closed-source; their pinned images (and the rtvi-vlm image tag) are listed in references/vss-source-layout.md.


RTVI-VLM: Two Deployment Methods

The implementation is selected by VLM_MODEL_TO_USE (standalone) or RTVI_VLM_MODEL_TO_USE (VSS blueprint).

Method A: OpenAI-Compatible Endpoint (openai-compat)

The RTVI container makes HTTP calls to an external inference server; it does not load the model itself. This covers local or remote NIM, external vLLM, and OpenAI.

Method B: vLLM Inside the RTVI Container (vllm-compatible)

The container downloads and serves the model using its bundled vLLM engine. No external inference server is needed.

Hard constraint: The model architecture must be supported by the vLLM version shipped in the selected RTVI image. Check the supported-model guidance in services/rtvi/rt-vlm/README.md. If the model requires a newer vLLM, use Method A with an external container instead.


Configuring RTVI-VLM via VSS Blueprint

The VSS blueprint .env / generated.env maps most RTVI_VLM_* variables to container-native names. The model or deployment identifier is the important exception: v3.2.1 maps VLM_NAME to VIA_VLM_OPENAI_MODEL_DEPLOYMENT_NAME. RTVI_VLM_PORT is also exceptional: it is required by Compose to publish the service and must be explicitly supplied when generated.env does not contain it.

Required port configuration — all Blueprint workflows

Add RTVI_VLM_PORT and update the two hardcoded 8018 lines in the Alerts profile .env, then rerun dev-profile.sh. Edit .env, not generated.env — the generator overwrites it.

bash
# ${VSS_PROFILE_DIR}/.env
RTVI_VLM_PORT=8018                                        # add this line
RTVI_VLM_BASE_URL=http://${HOST_IP}:${RTVI_VLM_PORT}     # update: was http://${HOST_IP}:8018
RTVI_VLM_ENDPOINT=http://${HOST_IP}:${RTVI_VLM_PORT}/v1  # update: was http://${HOST_IP}:8018/v1

Keep the value at 8018 unless that port is unavailable. Variable ownership, the VLM_BASE_URL generator overwrite, and the patch required for any other port are in references/port-and-url-wiring.md.

Method A — OpenAI-compatible endpoint

In the v3.2.1 blueprint Compose file, these host variables map as follows:

Blueprint variableContainer variable
RTVI_VLM_MODEL_TO_USEVLM_MODEL_TO_USE
RTVI_VLM_ENDPOINTVIA_VLM_ENDPOINT
RTVI_VLM_API_KEYVIA_VLM_API_KEY
VLM_NAMEVIA_VLM_OPENAI_MODEL_DEPLOYMENT_NAME
RTVI_VLM_MODEL_PATHMODEL_PATH

After applying the shared port configuration above, configure the active profile environment:

bash
# ${VSS_PROFILE_DIR}/.env
RTVI_VLM_MODEL_TO_USE=openai-compat
RTVI_VLM_ENDPOINT=http://host.docker.internal:30082/v1
VLM_NAME=<model-or-deployment-id>
OPENAI_API_KEY=<your-api-key>     # client fallback/default credential
RTVI_VLM_API_KEY=<your-api-key>   # mapped to preferred VIA_VLM_API_KEY
# RTVI_VLM_IMAGE_TAG=3.2.1    # pin to a specific image tag; defaults to 3.2.1

For an authenticated Method A endpoint, set OPENAI_API_KEY and RTVI_VLM_API_KEY to the same endpoint-specific credential. The stock Compose file maps RTVI_VLM_API_KEY to VIA_VLM_API_KEY, and the OpenAI-compatible client prefers VIA_VLM_API_KEY over OPENAI_API_KEY. Leave both unset only when the endpoint intentionally accepts unauthenticated requests.

Examples:

bash
# NIM on the Docker host. Stock Compose maps this name through host-gateway.
RTVI_VLM_ENDPOINT=http://host.docker.internal:30082/v1
VLM_NAME=nvidia/cosmos-reason2-8b

# NVIDIA API catalog (set keys via silent prompt)
RTVI_VLM_ENDPOINT=https://integrate.api.nvidia.com/v1
VLM_NAME=nvidia/cosmos-reason2-8b
OPENAI_API_KEY=<your-nvidia-api-key>
RTVI_VLM_API_KEY=<your-nvidia-api-key>

# OpenAI (set keys via silent prompt)
RTVI_VLM_ENDPOINT=https://api.openai.com/v1
VLM_NAME=gpt-4o
OPENAI_API_KEY=<your-openai-api-key>
RTVI_VLM_API_KEY=<your-openai-api-key>

Migrating from Method B: Replace RTVI_VLM_API_KEY before recreating rtvi-vlm; do not leave the previous NGC credential in that variable. If RTVI_VLM_API_KEY is empty, stock Compose falls back to NGC_CLI_API_KEY when populating VIA_VLM_API_KEY. Because the client prefers VIA_VLM_API_KEY, a stale NGC key would be sent to the new Method A endpoint instead of OPENAI_API_KEY. This is both an authentication failure and a credential-disclosure risk.

RTVI_VLM_OPENAI_MODEL_DEPLOYMENT_NAME is not consumed by the public v3.2.1 blueprint Compose file. Use VLM_NAME unless the deployed Compose file has been intentionally customized.

Verify the configured model identifier from inside the RTVI container. The conditional Authorization header keeps keyless local endpoints working while authenticating endpoints such as OpenAI and the NVIDIA API Catalog. Endpoints may advertise multiple models, so confirm the configured ID appears anywhere in the /models response:

bash
ALERTS_ENV=developer-profiles/dev-profile-alerts/generated.env

# Container-native name for `VLM_NAME`.
EXPECTED_MODEL="$(docker compose --env-file "${ALERTS_ENV}" exec -T rtvi-vlm \
  printenv VIA_VLM_OPENAI_MODEL_DEPLOYMENT_NAME | tr -d '\r')"
test -n "${EXPECTED_MODEL}"

docker compose --env-file "${ALERTS_ENV}" exec -T rtvi-vlm sh -lc \
  'set --
   if [ -n "${VIA_VLM_API_KEY:-}" ]; then
     set -- -H "Authorization: Bearer ${VIA_VLM_API_KEY}"
   fi
   curl -fsS "$@" "${VIA_VLM_ENDPOINT%/}/models"' \
  | jq -e --arg model "${EXPECTED_MODEL}" 'any(.data[]; .id == $model)' \
  || { echo "Endpoint does not advertise '${EXPECTED_MODEL}'" >&2; exit 1; }

jq -e exits non-zero when the assertion is false, so the check passes only when the configured model appears somewhere in the list. To see which models the endpoint offers, drop the jq filter and read the raw response.

Treat HTTP 401 from this request as an authentication failure. Only treat a successful response that fails the jq assertion as a model-name mismatch.

Show full SKILL.md (633 more words)Show less
Method B — vLLM inside RTVI container
bash
RTVI_VLM_MODEL_TO_USE=vllm-compatible
RTVI_VLM_MODEL_PATH=git:https://huggingface.co/Qwen/Qwen3-VL-30B-A3B-Instruct

# With HuggingFace token:
# HF_TOKEN=<your-hf-token>

# From NGC:
# RTVI_VLM_MODEL_PATH=ngc:nim/nvidia/cosmos-reason2-8b:hf-1208

MODEL_PATH prefix conventions (see public services/rtvi/rt-vlm/src/vlm_pipeline/ngc_model_downloader.py):

  • ngc:<registry>/<org>/<model>:<version> — download from NGC
  • git:<url> — clone from HuggingFace or any git LFS repo
  • /path/to/local/dir — use pre-downloaded local weights

A bad MODEL_PATH aborts startup during download, before the server binds its port, so the container exits non-zero. v3.2.1 raises a plain Exception: Failed to download model <name> from <url> for git: paths, and Model download failed with status code <code>, Could not authenticate with NGC., or Could not find the model. for ngc: paths.


Configuring RTVI-VLM Standalone

When running services/rtvi/rt-vlm/docker/compose.yaml from the public VSS checkout, use the container-native names:

bash
# services/rtvi/rt-vlm/docker/.env
BACKEND_PORT=8000

# Method A
VLM_MODEL_TO_USE=openai-compat
VIA_VLM_ENDPOINT=http://host.docker.internal:30082/v1
VIA_VLM_OPENAI_MODEL_DEPLOYMENT_NAME=nvidia/cosmos-reason2-8b
VIA_VLM_API_KEY=<your-api-key>       # if required; set via silent prompt

# Method B
VLM_MODEL_TO_USE=vllm-compatible
MODEL_PATH=git:https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct

The standalone Compose file defaults to nvcr.io/nvidia/vss-core/vss-rt-vlm:3.2.1 and requires BACKEND_PORT to be set. All supported VLM_MODEL_TO_USE values: openai-compat, vllm-compatible, cosmos-reason1, cosmos-reason2, cosmos-reason3, custom.

Both stock Compose definitions use bridged networking and map host.docker.internal through host-gateway. Choose the endpoint host by where the inference server runs:

  • Docker host: host.docker.internal
  • Same Compose network: the inference server's Compose service name
  • Remote machine: a routable hostname or IP address
  • localhost: only when the RTVI container explicitly uses host networking

For standalone Method A, verify connectivity from inside rtvi-server:

bash
docker compose exec -T rtvi-server sh -lc \
  'set --
   if [ -n "${VIA_VLM_API_KEY:-}" ]; then
     set -- -H "Authorization: Bearer ${VIA_VLM_API_KEY}"
   fi
   curl -fsS "$@" "${VIA_VLM_ENDPOINT%/}/models"'

See the environment-variable reference in public services/rtvi/rt-vlm/README.md.


Configuring vlm-as-verifier (2d_cv mode only)

vlm-as-verifier runs as the alert-bridge service and is separate from RTVI-VLM. Its config is bind-mounted from ${VSS_PROFILE_DIR}/vlm-as-verifier/configs/config.yml.

Keep the config parameterized:

yaml
# ${VSS_PROFILE_DIR}/vlm-as-verifier/configs/config.yml
vlm:
  base_url: ${VLM_BASE_URL}/v1
  model: ${VLM_NAME}
  max_tokens: 4096

For the v3.2.1 local-VLM Alerts flow, dev-profile.sh --vlm-device-id ... populates VLM_BASE_URL with http://<host-ip>:8018 and sets VLM_NAME in the generated environment. Remote mode likewise writes the selected endpoint. Only hardcode these fields when deliberately overriding the profile-generated values.

The 8018 here is literal in the generator, not derived from RTVI_VLM_PORT. If RTVI-VLM is published on any other port, alert-bridge still dials 8018 until you patch VLM_BASE_URL in generated.env (references/port-and-url-wiring.md).

After editing the config or environment, recreate the service:

bash
cd "${VSS_DEPLOY_DIR}"
docker compose \
  --env-file developer-profiles/dev-profile-alerts/generated.env \
  up -d --force-recreate alert-bridge

Configuring vss-agent

vss-agent is another separate consumer. The v3.2.1 Alerts profile selects a block in ${VSS_PROFILE_DIR}/vss-agent/configs/config.yml using VLM_MODEL_TYPE. The local RTVI-VLM flow uses rtvi:

yaml
rtvi_vlm:  # VLM_MODEL_TYPE=rtvi
  base_url: ${RTVI_VLM_BASE_URL}/v1
nim_vlm:   # VLM_MODEL_TYPE=nim
  base_url: ${VLM_BASE_URL}/v1
openai_vlm: # VLM_MODEL_TYPE=openai
  base_url: ${VLM_BASE_URL}/v1
vllm_vlm:  # VLM_MODEL_TYPE=vllm
  base_url: ${VLM_BASE_URL}/v1

For a v3.2.1 local Alerts deployment, keep the generated values, which are normally equivalent to:

bash
VLM_MODEL_TYPE=rtvi
VLM_NAME=<model-id>
RTVI_VLM_PORT=8018                                    # from .env
RTVI_VLM_BASE_URL=http://<host-ip>:${RTVI_VLM_PORT}   # from .env
VLM_BASE_URL=http://<host-ip>:8018                    # written by dev-profile.sh

RTVI_VLM_PORT is required by the rtvi-vlm Compose port mapping and has no Compose default. Define it in the Alerts profile .env and regenerate generated.env before using any of these Compose commands. The two URLs have different owners and agree only at the default port; see references/port-and-url-wiring.md.

For a remote NIM, OpenAI-compatible endpoint, or external vLLM, select the matching nim, openai, or vllm profile and set VLM_BASE_URL without a trailing /v1 because the config appends it. Recreate vss-agent after changing these values.


Health Checks

After changing VLM configuration, recreate the affected service and inspect it through the same Alerts Compose model:

bash
cd "${VSS_DEPLOY_DIR}"
ALERTS_ENV=developer-profiles/dev-profile-alerts/generated.env
# Required when an existing `generated.env` does not yet contain the port.
# The persistent fix is to add it to the profile `.env` and regenerate.
export RTVI_VLM_PORT=8018

# Recreate RTVI-VLM and wait up to 10 minutes for its Compose healthcheck.
docker compose --env-file "${ALERTS_ENV}" \
  up -d --force-recreate --wait --wait-timeout 600 rtvi-vlm

# The container must still be running. A failed weights download exits here.
CID="$(docker compose --env-file "${ALERTS_ENV}" ps -aq rtvi-vlm)"
test "$(docker inspect -f '{{.State.Status}}' "${CID}")" = running || {
  docker logs --tail 50 "${CID}" >&2; exit 1; }

# Require the absence of a download or warm-up failure. `Warmup ... done` is
# logged on both paths in v3.2.1, so only the error lines are conclusive.
docker logs "${CID}" 2>&1 | grep -Eq \
  'Failed to download model|Model download failed|Could not find the model|Could not authenticate with NGC|Error during (model )?warmup' \
  && { echo 'rtvi-vlm failed to download weights or warm up' >&2; exit 1; }

# Confirm the readiness endpoint from inside the container.
docker compose --env-file "${ALERTS_ENV}" exec -T rtvi-vlm \
  curl -fsS http://localhost:8000/v1/health/ready

# Capture the implementation and the model ID advertised by the live service.
VLM_METHOD="$(docker compose --env-file "${ALERTS_ENV}" exec -T rtvi-vlm \
  printenv VLM_MODEL_TO_USE | tr -d '\r')"
ACTUAL_MODEL="$(docker compose --env-file "${ALERTS_ENV}" exec -T rtvi-vlm \
  curl -fsS http://localhost:8000/v1/models | jq -er '.data[0].id')"
test -n "${ACTUAL_MODEL}"

# Readiness and /v1/models only report metadata. Require one real inference; a
# text-only chat request needs no media asset and works on both methods.
docker compose --env-file "${ALERTS_ENV}" exec -T rtvi-vlm \
  curl -fsS --max-time 120 -X POST http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d "$(jq -nc --arg m "${ACTUAL_MODEL}" '{model: $m, max_tokens: 16, stream: false,
        messages: [{role: "user", content: "Reply with: ok"}]}')" \
  | jq -e '((.choices[0].message.content // "") | length) > 0' > /dev/null \
  || { echo 'inference smoke test returned no content' >&2; exit 1; }

Readiness is built from is_alive() per process, so it reports healthy: true for a live process whose model never warmed up, and /v1/chat/completions returns HTTP 200 with empty content when the pipeline produced no output. The jq assertion on non-empty content is what makes the last check meaningful.

Model identity differs by method

The two methods advertise their model ID differently, so the follow-up check is not the same:

  • Method A (openai-compat) — verify that the configured VIA_VLM_OPENAI_MODEL_DEPLOYMENT_NAME appears anywhere in the /v1/models list using the any(.data[]; .id == $model) check from the Method A endpoint verification section. Treat HTTP 401 as an authentication failure; treat a 200 response where the model is absent as a model-name mismatch.
  • Method B (vllm-compatible) — do not compare the advertised ID against VLM_NAME or VIA_VLM_OPENAI_MODEL_DEPLOYMENT_NAME. The resolved model directory basename becomes the advertised /v1/models ID, so ngc:nim/nvidia/cosmos-reason2-8b:hf-1208 is served as nim_nvidia_cosmos-reason2-8b_hf-1208. That difference is expected for Method B and does not indicate a deployment failure. Point downstream consumers (vlm-as-verifier, vss-agent) at that advertised ID, never at the NGC or Hugging Face path.

Run the per-method verification commands from references/health-checks.md — covers the Method A identity check, the Method B inspection, warm-up diagnostics, and querying a specific external endpoint.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/rtvi-vlm-customize-model of NVIDIA/skills.

  • SKILL.md
  • .gitkeep
  • BENCHMARK.md
  • evals/evals.json
  • references/health-checks.md
  • references/port-and-url-wiring.md
  • references/vss-source-layout.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Rtvi Vlm Customize Model next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rtvi Vlm Customize Model compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rtvi Vlm Customize Model this skillNVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Vllm Deploy K8svllm-project/vllm-skills103—~2kAutomated safety check: PassApache-2.0
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Vllm Deploy Simplevllm-project/vllm-skills103—~1.6kAutomated safety check: PassApache-2.0
LLM Gatewaysickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy K8s

    vllm-project/vllm-skills

    Deploy vLLM to Kubernetes (K8s) with GPU support, health probes, and OpenAI-compatible API endpoint.

    103 GitHub stars~2k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Vllm Deploy Simple

    vllm-project/vllm-skills

    Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Gateway

    sickn33/agentic-awesome-skills

    Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Backend & APIsAuto-check passed
  • Ray LLM

    pproenca/dot-skills

    LLM workloads on open-source Ray (pinned to 2.57) — OpenAI-compatible serving with ray.serve.llm (vLLM-backed LLMConfig + buildopenaiapp) and batch inference with ray.data.llm (buildprocessor).

    214 GitHub stars~1.7k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Rtvi Vlm Customize Model

What does Rtvi Vlm Customize Model do?

How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks. Rtvi Vlm Customize Model is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. How to swap the VLM in the VSS Alerts Blueprint — covers RTVI-VLM microservice deployment methods, all three VLM consumers (rtvi-vlm, vlm-as-verifier, vss-agent), and health checks.

When should I use Rtvi Vlm Customize Model?

Rtvi Vlm Customize Model fits situations like: tasks that involve Microservices; tasks that involve Deployment; tasks that involve LLM inference and serving.

How do I install Rtvi Vlm Customize Model in Claude Code?

Run `npx skills add NVIDIA/skills --skill rtvi-vlm-customize-model -a claude-code`. Or copy the skill folder (skills/rtvi-vlm-customize-model in NVIDIA/skills) into .claude/skills/rtvi-vlm-customize-model in your project. Claude Code loads it when a task matches its description.

How do I install Rtvi Vlm Customize Model in Codex?

Run `npx skills add NVIDIA/skills --skill rtvi-vlm-customize-model -a codex`. Or copy the skill folder (skills/rtvi-vlm-customize-model in NVIDIA/skills) into .agents/skills/rtvi-vlm-customize-model in your project. Codex loads it when a task matches its description.

Can I use Rtvi Vlm Customize Model in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill rtvi-vlm-customize-model -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rtvi-vlm-customize-model, .gemini/skills/rtvi-vlm-customize-model, .github/skills/rtvi-vlm-customize-model and .opencode/skills/rtvi-vlm-customize-model in your project.

What does Rtvi Vlm Customize Model need to run?

Going by SKILL.md and its folder, Rtvi Vlm Customize Model needs the command-line tools its instructions call (docker, curl and jq) and credentials named VIA_VLM_API_KEY, RTVI_VLM_API_KEY, OPENAI_API_KEY and HF_TOKEN. Our summary lists: Docker; A credential in RTVI_VLM_API_KEY; A credential in VIA_VLM_API_KEY.

Does Rtvi Vlm Customize Model access the network?

SKILL.md names 4 domains. In commands or code: huggingface.co, integrate.api.nvidia.com and api.openai.com; the agent is likely to contact these when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.

Is Rtvi Vlm Customize Model safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Rtvi Vlm Customize Model use?

Rtvi Vlm Customize Model is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rtvi Vlm Customize Model use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Rtvi Vlm Customize Model?

Skills that share tags, products or a category with Rtvi Vlm Customize Model: vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars), Vllm Deploy K8s (vllm-project/vllm-skills, 103 stars), Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars) and Vllm Deploy Simple (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rtvi Vlm Customize Model?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.