A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Rtvi Byom Porting

skills CLI
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-byom-porting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization rtvi-byom-porting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/deployment/rtvi-byom-porting .claude/skills/rtvi-byom-porting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rtvi-byom-porting
GitHub stars
1.9k
Token cost
~1.4k tokens
SKILL.md length
635 words
Files
6 (incl. scripts, references)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

  • Works in 4 steps: Configuration only: use… → RTVI adapter: normalize config,… → Plugin or shim: register architecture… → …
  • Validating a bring-your-own VLM in VSS RT-VLM
  • SKILL.md covers Purpose, Model Contract, Integration Decision and Hard Gates, plus 2 more sections
  • Runs Python scripts from its folder; calls python3; needs HF_TOKEN

What it does

Rtvi Byom Porting is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Use when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and model-specific runtime dependencies. Not for selecting an already-supported model or ordinary RT-VLM deployment.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `agents/openai.yaml`, `references/quality-gates.md` and `references/vllm-porting.md`).

It sits in AI & LLM Engineering, covering LLM inference and serving, Model hubs and datasets and Deployment. It works with vLLM and Hugging Face. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.

When your agent uses it

  • Validating a bring-your-own VLM in VSS RT-VLM
  • Including custom Hugging Face
  • NGC checkpoints
  • Model-specific runtime dependencies

Example prompts

  • “/rtvi-byom-porting”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Configuration only: use VLM_MODEL_TO_USE=vllm-compatible and
  2. RTVI adapter: normalize config, processor, request or response behavior
  3. Plugin or shim: register architecture and deterministic weight mappings
  4. Version-gated patch: patch vLLM only after recording why the first three

What it can do on your machine

Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rtvi Byom Porting loads about 1.4k tokens when it runs, and up to ~3.2k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 635 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 635 words, ~1,398 tokens.

Download SKILL.mdSave it as .claude/skills/rtvi-byom-porting/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
rtvi-byom-porting
description
Use when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and model-specific runtime dependencies. Not for selecting an already-supported model or ordinary RT-VLM deployment.
license
Apache-2.0
metadata.version
3.3.0-rc0
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
metadata.tags
nvidia blueprint rtvi vlm byom vllm model-porting

RTVI BYOM Porting

Purpose

Port a VLM into the VSS RT-VLM service without hard-coding one checkpoint or GPU platform. Prefer configuration, then an RTVI adapter or vLLM plugin, and patch vLLM only when the supported extension points cannot express the model.

Use vss-deploy-dense-captioning for ordinary deployment or an already-supported model. Use this skill when model architecture, processor, weight mapping, runtime dependencies, or request adaptation requires repository work.

Model Contract

Before changing code, record:

  • immutable model source and revision, license and access requirements;
  • architecture, processor, tokenizer, quantization and context length;
  • vision-token and frame-sampling assumptions;
  • CUDA architecture, memory, vLLM/runtime and custom-kernel requirements.

For a private Hugging Face source, use HF_TOKEN only during authenticated download/cache population. Never print it, embed it in MODEL_PATH, or bake it into an image. Prove the cached model starts after the temporary credential is removed. If remote model code is required, review it first and pair VLM_TRUST_REMOTE_CODE=true with an exact allowlist entry. For the VSS profile, set RTVI_VLM_MODEL_PATH_ALLOWLIST; inside the service container this becomes RTVI_MODEL_PATH_ALLOWLIST. Likewise, the profile input RTVI_VLM_ALLOW_UNSAFE_MODEL_CONFIG becomes RTVI_ALLOW_UNSAFE_MODEL_CONFIG; do not enable it without reviewing the blocked config hooks.

Integration Decision

Stop at the first path that works:

  1. Configuration only: use VLM_MODEL_TO_USE=vllm-compatible and MODEL_PATH=<ngc: or mounted path>. A git: source is acceptable only for exploration because the develop downloader does not pin Hugging Face revisions; use a revision-pinned mounted snapshot for reproducible evidence.
  2. RTVI adapter: normalize config, processor, request or response behavior inside services/rtvi/rt-vlm/ while preserving existing model behavior.
  3. Plugin or shim: register architecture and deterministic weight mappings without editing vendored vLLM internals.
  4. Version-gated patch: patch vLLM only after recording why the first three paths cannot work. Keep it optional and add a focused regression test.

Use VLM_MODEL_TO_USE=custom with MODEL_IMPLEMENTATION_PATH only for a custom RTVI model implementation. In the VSS Compose profile these are exposed as RTVI_VLM_MODEL_TO_USE, RTVI_VLM_MODEL_PATH, and RTVI_VLM_MODEL_IMPLEMENTATION_PATH.

For source, adapter, plugin, or custom-backend changes, build the RT-VLM image from services/rtvi/rt-vlm/, test it with standalone Compose via RTVI_IMAGE, then select the same repository and tag in the VSS profile via VSS_RT_VLM_IMAGE and VSS_RT_VLM_TAG. Record the tested image digest. A host MODEL_IMPLEMENTATION_PATH alone is insufficient: the implementation must exist at that path inside the selected image or an explicit Compose bind mount.

Read references/vllm-porting.md before changing model loading, registration, weight mapping, kernels, cache behavior or vLLM.

Show full SKILL.md (249 more words)Show less

Hard Gates

  • Do not enable eager mode by default. If it is unavoidable, capture the exact blocker, performance impact and removal condition.
  • Do not bind generic model logic to H100, B200, RTX, Orin, x86 or Jetson. Capability-detect or use existing configuration boundaries.
  • Keep secrets out of commands, logs, images, reports and committed files.
  • Do not accept a successful load as proof of a successful port. Output must be legible and grounded for every claimed modality.

Validation

Run the smallest relevant checks in this order:

  1. Static checks and focused tests for changed Python, shell and configuration.
  2. Start RT-VLM through the canonical VSS Compose/profile path and verify /v1/health/ready and /v1/models. First use standalone Compose when a revision-pinned host model snapshot must be mounted; the VSS profile does not expose MODEL_ROOT_DIR on develop.
  3. Smoke-test text, image and video inputs for every claimed modality. Set the explicit media type when testing images rather than relying on video routing.
  4. Read references/quality-gates.md, retain raw redacted responses, and reject gibberish, repetition or ungrounded claims.
  5. Run relevant caption, event-detection or summarization accuracy checks.
  6. Record latency, throughput, GPU memory/utilization, eager state and custom kernel state.

Use skills/deployment/rtvi-byom-porting/scripts/byom_port_report.py to turn observed JSON facts into a compact Markdown report. Evidence is mandatory for a PASS result; the helper does not manufacture it. Run python3 skills/deployment/rtvi-byom-porting/scripts/tests/test_byom_port_report.py after changing the helper.

Completion Contract

Report the exact model revision and backend, integration path, eager-mode state, platform gating, smoke/accuracy/performance evidence, remaining blockers and the next bounded validation step.

© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/deployment/rtvi-byom-porting of NVIDIA-AI-Blueprints/video-search-and-summarization.

  • SKILL.md
  • agents/openai.yaml
  • references/quality-gates.md
  • references/vllm-porting.md
  • scripts/byom_port_report.py
  • scripts/tests/test_byom_port_report.py

Open the folder on GitHubat commit fdb6a7a

Compare with similar skills

Rtvi Byom Porting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rtvi Byom Porting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rtvi Byom Porting this skillNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.4kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hf Cloud Serving Image Selectionwaybarrios/opencode-power-pack534—~4.3kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Add Modelguoqingbao/xinfer334—~4.2kAutomated safety check: NotesMIT
Resolvealexziskind1/model-shelf130—~792Automated safety check: PassMIT

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Hf Cloud Serving Image Selection

    waybarrios/opencode-power-pack

    Select and verify the current region-specific serving container URI for a SageMaker model deployment.

    534 GitHub stars~4.3k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Model

    guoqingbao/xinfer

    Adapt and port new LLM model architectures to this xinfer project.

    334 GitHub stars~4.2k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Resolve

    alexziskind1/model-shelf

    Always resolve Hugging Face models via model-shelf before any download.

    130 GitHub stars~792 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Check Model

    guoqingbao/xinfer

    Check model compatibility with xinfer before loading. An agent skill from guoqingbao/xinfer.

    334 GitHub stars~3.8k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA-AI-Blueprints/video-search-and-summarization

All 22 skills in this repo
  • Benchmark Video Search

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

    1.9k GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…

    1.9k GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Rtvi Vlm Perf Testing

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.

    1.9k GitHub stars~8.6k tokensUpdated today
    Auto-check: notes
  • Vss Build Vision AI

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…

    1.9k GitHub stars~15k tokensUpdated today
    Auto-check: notes
  • Vss Evaluate Caption Accuracy

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

    1.9k GitHub stars~2.1k tokensUpdated today
    Auto-check: notes
  • Vss Manage Alerts

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when operating VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, verification verdicts, on-demand verification, always-on operation, Slack…

    1.9k GitHub stars~11k tokensUpdated today
    Auto-check: warnings

Questions about Rtvi Byom Porting

What does Rtvi Byom Porting do?

A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…. Rtvi Byom Porting is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Use when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and model-specific runtime dependencies.

When should I use Rtvi Byom Porting?

Rtvi Byom Porting fits situations like: validating a bring-your-own VLM in VSS RT-VLM; including custom Hugging Face; NGC checkpoints; model-specific runtime dependencies.

How do I install Rtvi Byom Porting in Claude Code?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-byom-porting -a claude-code`. Or copy the skill folder (skills/deployment/rtvi-byom-porting in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/rtvi-byom-porting in your project. Claude Code loads it when a task matches its description.

How do I install Rtvi Byom Porting in Codex?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-byom-porting -a codex`. Or copy the skill folder (skills/deployment/rtvi-byom-porting in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/rtvi-byom-porting in your project. Codex loads it when a task matches its description.

Can I use Rtvi Byom Porting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill rtvi-byom-porting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rtvi-byom-porting, .gemini/skills/rtvi-byom-porting, .github/skills/rtvi-byom-porting and .opencode/skills/rtvi-byom-porting in your project.

What does Rtvi Byom Porting need to run?

Going by SKILL.md and its folder, Rtvi Byom Porting needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named HF_TOKEN. Our summary lists: Python 3.

Does Rtvi Byom Porting access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rtvi Byom Porting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Rtvi Byom Porting use?

Rtvi Byom Porting is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rtvi Byom Porting use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.8k tokens, read only when the agent opens those files.

What are the alternatives to Rtvi Byom Porting?

Skills that share tags, products or a category with Rtvi Byom Porting: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hf Cloud Serving Image Selection (waybarrios/opencode-power-pack, 534 stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Add Model (guoqingbao/xinfer, 334 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rtvi Byom Porting?

NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.

Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.