Official agent skill

Tao Finetune Video Clip

by NVIDIA in NVIDIA/skills

InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Finetune Video Clip

skills CLI
$ npx skills add NVIDIA/skills --skill tao-finetune-video-clip -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-finetune-video-clip --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-finetune-video-clip .claude/skills/tao-finetune-video-clip && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-finetune-video-clip
GitHub stars
3.6k
Token cost
~3.5k tokens
SKILL.md length
966 words
Files
17 (incl. references)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment.

  • The user asks to fine-tune IV2CLIP
  • SKILL.md covers Train Action Policy, Quick Start (local Docker), TensorRT Deploy and Quick Start (virtualenv —…, plus 8 more sections
  • Calls docker and python; needs NGC_KEY and HF_TOKEN
  • Run videoclip train/evaluate/inference/export

What it does

Tao Finetune Video Clip is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment. Use when the user asks to "fine-tune IV2CLIP", "run videoclip train/evaluate/inference/export", "build a Video-CLIP TensorRT engine", "InternVideo2-CLIP on KPI chunks", or "TAO videoclip on vadr1chunks JSON".

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit and the pinned TAO videoclip PyTorch and Deploy containers (see references/skillinfo.yaml and…

It sits in AI & LLM Engineering, covering Video production, Fine-tuning and LLM inference and serving. It works with NVIDIA AI Platform, ONNX and PyTorch. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user asks to fine-tune IV2CLIP
  • Run videoclip train/evaluate/inference/export
  • Build a Video-CLIP TensorRT engine
  • InternVideo2-CLIP on KPI chunks

Example prompts

  • “fine-tune IV2CLIP”
  • “run videoclip train/evaluate/inference/export”
  • “build a Video-CLIP TensorRT engine”
  • “/tao-finetune-video-clip”

Requirements

  • Python 3
  • Docker
  • A credential in NGC_KEY
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit and the pinned TAO video_clip PyTorch and Deploy containers (see references/skill_info.yaml and references/tao-deploy-video-clip.skill_info.yaml), or a local tao-pytorch checkout + tao-cli venv for PyTorch virtualenv runs. MobileCLIP + InternVideo2 weights must be on disk for offline eval (HF LFS may be blocked in CI). Metadata JSON uses vadr1_chunks with absolute video_path entries.
  • Pre-approved tools (allowed-tools): Read, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NGC_KEY
    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit and the pinned TAO video_clip PyTorch and Deploy containers (see references/skill_info.yaml and references/tao-deploy-video-clip.skill_info.yaml), or a local tao-pytorch checkout + tao-cli venv for PyTorch virtualenv runs. MobileCLIP + InternVideo2 weights must be on disk for offline eval (HF LFS may be blocked in CI). Metadata JSON uses vadr1_chunks with absolute video_path entries.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Finetune Video Clip loads about 3.5k tokens when it runs, and up to ~7.4k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 966 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 966 words, ~3,453 tokens.

Download SKILL.mdSave it as .claude/skills/tao-finetune-video-clip/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.
name
tao-finetune-video-clip
description
InternVideo2-CLIP L14 (TAO video_clip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment. Use when the user asks to "fine-tune IV2CLIP", "run video_clip train/evaluate/inference/export", "build a Video-CLIP TensorRT engine", "InternVideo2-CLIP on KPI chunks", or "TAO video_clip on vadr1_chunks JSON".
allowed-tools
Read, Bash
compatibility
Requires docker + nvidia-container-toolkit and the pinned TAO video_clip PyTorch and Deploy containers (see references/skill_info.yaml and references/tao-deploy-video-clip.skill_info.yaml), or a local tao-pytorch checkout + tao-cli venv for PyTorch virtualenv runs. MobileCLIP + InternVideo2 weights must be on disk for offline eval (HF LFS may be blocked in CI). Metadata JSON uses vadr1_chunks with absolute video_path entries.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
video, vision-language, vlm, multimodal, retrieval, embedding, internvideo2, iv2clip, fine-tuning, deployment

InternVideo2-CLIP (TAO video_clip)

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

TAO task video_clip wraps OpenGVLab InternVideo2-CLIP L14. The PyTorch image provides train, evaluate, inference, export, and default_specs. TAO Deploy provides gen_trt_engine, TensorRT evaluate, and TensorRT inference.

Container images and per-action commands are in references/skill_info.yaml and references/tao-deploy-video-clip.skill_info.yaml. Starting specs are in references/spec_template_*.yaml.

Release note: The pinned PyTorch image is the TAO 7.2 release-candidate build validated for Video-CLIP. It includes PyAV 17.1.0 as the primary decoder and ONNXScript 0.7.1 for export, with decord absent. The TAO Deploy image is pinned independently because gen_trt_engine and TensorRT-backed actions do not run in the PyTorch image.

Known-broken images: interim builds cut before tao-pytorch commit 0cc31de4 ship a video_clip package with no model.backbones submodule, so train/evaluate/inference die at import while video_clip --help still exits 0. Images without PyAV also fail at data loading. Run both import checks in the preflight below before pulling data or launching a run.

Train Action Policy

AutoML is not packaged for this model skill. Always use direct video_clip actions even when a higher-level request mentions AutoML. Non-train actions stay in this skill.

Quick Start (local Docker)

Use the pinned TAO container declared in references/skill_info.yaml. Pull with NGC_KEY when the image is not cached locally.

bash
VIDEO_CLIP_IMAGE_DEFAULT="nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt"  # versions-key: images.tao_toolkit.pyt
VIDEO_CLIP_IMAGE="${VIDEO_CLIP_IMAGE:-$VIDEO_CLIP_IMAGE_DEFAULT}"
docker pull "$VIDEO_CLIP_IMAGE"

Expected workspace layout (host paths bind-mounted into the container):

text
workspace/
├── data/
│   ├── train.json              # vadr1_chunks metadata; video_path = /data/videos/<name>.mp4
│   ├── val.json
│   ├── prompts.txt             # one text prompt per line (inference)
│   └── videos/                 # mp4 clips referenced by train.json / val.json
├── model/
│   ├── mobileclip_blt.pt       # MobileCLIP weights for model.text_encoder
│   └── hf/                     # offline InternVideo2_distillation_models snapshot (HF_HUB_OFFLINE=1)
├── specs/
│   ├── train.yaml
│   ├── evaluate.yaml
│   ├── inference.yaml
│   └── export.yaml
├── deploy_specs/
│   ├── gen_trt_engine.yaml
│   ├── evaluate.yaml
│   └── inference.yaml
└── results/

Docker options for all actions (skill-eval CI uses the same $WORKSPACE_DIR bind-mount pattern):

bash
VIDEO_CLIP_IMAGE_DEFAULT="nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt"  # versions-key: images.tao_toolkit.pyt
VIDEO_CLIP_IMAGE="${VIDEO_CLIP_IMAGE:-$VIDEO_CLIP_IMAGE_DEFAULT}"
RUN_ROOT="${RUN_ROOT:-$PWD}"
DOCKER_COMMON=(
  --rm --gpus all --shm-size=8g --network=host
  --shm-size=64g
  --ulimit memlock=-1
  --ulimit stack=67108864
  -e WANDB_DISABLED=true
  -e WANDB_MODE=disabled
  -e HF_HUB_OFFLINE=1
  -e HUGGINGFACE_HUB_CACHE=/model/hf
  -e TRANSFORMERS_OFFLINE=1
  -e TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1
  -v "$RUN_ROOT/data:/data:ro"
  -v "$RUN_ROOT/model:/model:ro"
  -v "$RUN_ROOT/specs:/specs:ro"
  -v "$RUN_ROOT/results:/results"
)

Preflight (host):

bash
[ -f "$RUN_ROOT/model/mobileclip_blt.pt" ] || echo "MISSING: MobileCLIP weights"
[ -f "$RUN_ROOT/data/train.json" ] || echo "MISSING: train metadata"
[ -d "$RUN_ROOT/model/hf" ] || echo "MISSING: offline HF snapshot under model/hf"
docker run --rm "$VIDEO_CLIP_IMAGE" video_clip --help >/dev/null || echo "MISSING: video_clip in container"
docker run --rm "$VIDEO_CLIP_IMAGE" \
  python -c "import nvidia_tao_pytorch.multimodal.video_clip.model.adapters.internvideo2clip" \
  >/dev/null 2>&1 || echo "BROKEN IMAGE: video_clip package is incomplete (missing model.backbones) — stop, see Release note"
docker run --rm "$VIDEO_CLIP_IMAGE" python -c "import av; print(av.__version__)" \
  >/dev/null 2>&1 || echo "BROKEN IMAGE: PyAV is missing — use the pinned FC image; do not install decord"
nvidia-smi >/dev/null 2>&1 || echo "note: no GPU visible"

Train:

bash
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
  video_clip train -e /specs/train.yaml results_dir=/results

Evaluate:

bash
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
  video_clip evaluate -e /specs/evaluate.yaml results_dir=/results

Inference:

bash
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
  video_clip inference -e /specs/inference.yaml results_dir=/results

Export:

bash
docker run "${DOCKER_COMMON[@]}" "$VIDEO_CLIP_IMAGE" \
  video_clip export -e /specs/export.yaml results_dir=/results

TensorRT Deploy

Use the independently pinned TAO Deploy image after PyTorch export. Read references/tao-deploy-video-clip.md before running the deploy actions; its templates cover the engine build, retrieval evaluation, and embedding inference contracts.

bash
VIDEO_CLIP_DEPLOY_IMAGE_DEFAULT="nvcr.io/nvidia/tao/tao-toolkit:7.2.0-deploy"  # versions-key: images.tao_toolkit.deploy
VIDEO_CLIP_DEPLOY_IMAGE="${VIDEO_CLIP_DEPLOY_IMAGE:-$VIDEO_CLIP_DEPLOY_IMAGE_DEFAULT}"

# Verify the independently pinned image before staging artifacts or using a GPU.
docker run --rm "$VIDEO_CLIP_DEPLOY_IMAGE" video_clip gen_trt_engine --help >/dev/null || \
  { echo "BROKEN IMAGE: Video-CLIP deploy entrypoint is unavailable" >&2; exit 1; }

docker run --gpus all --rm --shm-size=16g \
  -v "$RUN_ROOT/deploy_specs:/specs:ro" \
  -v "$RUN_ROOT/results/export:/models:ro" \
  -v "$RUN_ROOT/data:/data:ro" \
  -v "$RUN_ROOT/results/deploy:/results" \
  "$VIDEO_CLIP_DEPLOY_IMAGE" \
  video_clip gen_trt_engine -e /specs/gen_trt_engine.yaml

Keep the exported ONNX file, its matching *_config.yaml, and its matching *_tokenizer/ directory together. gen_trt_engine copies the sidecars beside the engine so TensorRT evaluate and inference can reconstruct preprocessing and tokenization.

Quick Start (virtualenv — local dev hosts)

On hosts with a tao-pytorch checkout and tao-cli venv (for example rtdetr-pytorch), run through tao-run-on-virtualenv instead of Docker:

bash
export VENV="${VENV:?set to your tao-cli virtualenv (must contain bin/video_clip)}"
export TAO_PYTORCH_ROOT="${TAO_PYTORCH_ROOT:?set to your tao-pytorch checkout with multimodal/video_clip}"
export PATH="$VENV/bin:$PATH"
export PYTHONPATH="$TAO_PYTORCH_ROOT:$TAO_PYTORCH_ROOT/tao-core:$PYTHONPATH"
export HF_HOME="${HF_HOME:-$PWD/hf_cache}"
export HF_HUB_OFFLINE="${HF_HUB_OFFLINE:-1}"
export WANDB_DISABLED=true
export WANDB_MODE=disabled

Copy a template spec from references/spec_template_*.yaml, fill checkpoint paths and metadata, then run:

bash
video_clip train -e /path/to/train.yaml
video_clip evaluate -e /path/to/evaluate.yaml
video_clip inference -e /path/to/inference.yaml
video_clip export -e /path/to/export.yaml

Credentials

  • NGC_KEY — pull the pinned TAO container from nvcr.io when it is not cached locally.
  • HF_TOKEN (online Hugging Face download only): Hugging Face read token used when model.vision_encoder and model.clip_head are null and the resolver downloads the InternVideo2 snapshot named by model.internvideo2clip_hf_id. For offline CI/eval, stage the complete S3/local snapshot and point both fields at the local files; no Hugging Face token is then required.

Treat tokens as secrets. Export them into the environment or pass them through an --env-file of bare KEY=value lines, rather than inlining values into generated spec YAML, command lines, or anything written under results_dir.

Data format (vadr1_chunks)

Metadata is a top-level JSON list of video records. Each record has video_path, split, and nested chunks[] with caption fields (queries, action_queries, anomaly_queries, dense_caption, scene_caption).

Point dataset.*.video_text.metadata at the user's JSON files (for example /data/train.json and /data/val.json in the templates). Each video_path must be an absolute path resolvable inside the runtime — for Docker, remap host clips to container paths like /data/videos/<name>.mp4 and set data_root: null unless using path_prefix_mapping.

Smoke overrides

For a short functional check (for example 2 epochs, 1 GPU, small batch):

yaml
train.num_epochs: 2
train.num_gpus: 1
train.gpu_ids: [0]
train.optim.warmup_steps: 10
dataset.train.batch_size: 2
dataset.val.batch_size: 2
dataset.train.num_workers: 4
dataset.val.num_workers: 4
dataset.train.video_text.caption_fields: [queries, action_queries, anomaly_queries]
dataset.train.video_text.caption_mode: first
dataset.metrics.mode: classification

Use dataset.metrics.mode: retrieval only when dataset.val.video_text.relevance_file is provided.

Inference

  • inference.mode: embeddings writes video_embeddings.h5 and text_embeddings.h5 under results_dir.
  • Provide inference.query.text_file (one prompt per line) and/or inference.query.input_texts.
  • Gallery videos come from dataset.inference.video_text.metadata.
Show full SKILL.md (404 more words)Show less

Export

export.encoder_type: combined produces the image-and-text ONNX consumed by the Video-CLIP deploy workflow. Keep export.batch_size: -1 for symbolic/dynamic batch dimensions; a positive value produces a fixed-batch ONNX. Export also writes matching *_config.yaml and *_tokenizer/ sidecars; preserve all three artifacts. Default opset is 23 on the vendor branch. Export requires a trained .pth at export.checkpoint.

LoRA

For vision-LoRA runs, start from tao-pytorch experiment_spec_lora.yaml or add a top-level peft: block (see shipped spec comments). Merge LoRA before export when checkpoints contain lora_* keys.

Common pitfalls

  • PATH must prefer $VENV/bin on virtualenv hosts so child processes resolve the venv Python.
  • evaluate uses dataset.val, not a separate test split — Lightning stage "test" still loads val metadata.
  • Eval precision: match train.precision (typically bf16) or flash-attn paths may fail under fp32 eval.
  • Empty action_queries on normal chunks become literal "Normal" positives during training; exclude Normal/Abnormal in dataset.metrics.exclude_categories for classification eval.
  • No hard-negative / explicit-neg training on the vendor branch unless the spec and branch explicitly enable it.
  • video_clip --help is not a health check. It exits 0 on an image whose video_clip package is missing model.backbones; only the import smoke check in the preflight catches it.
  • PyAV is the primary Video-CLIP decoder in TAO 7.2. The loader can fall back to the image's FFmpeg CLI and OpenCV support. A missing av import is an image defect: use the pinned FC image. Do not add or force-install decord, because it is not part of the supported TAO 7.2 decode contract.
  • TensorRT actions use TAO Deploy. gen_trt_engine, TensorRT evaluate, and TensorRT inference must use the independently pinned deploy image and deploy templates, not the PyTorch image/specs.
  • Deploy sidecars are required. TensorRT evaluation and text inference need the exported *_config.yaml and *_tokenizer/ beside the engine. Keep them with the ONNX input so engine generation can copy them automatically.
  • PyTorch ≥ 2.6 defaults to torch.load(weights_only=True) and rejects the TAO checkpoint’s numpy dtype objects with _pickle.UnpicklingError. TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1 is set in DOCKER_COMMON above; keep it for evaluate, inference, and export.
  • model/hf/ is a snapshot, not an HF hub cache. Offline packs stage InternVideo2 weights at repo-relative paths (stage1/L14/L14_dist_1B_stage2/pytorch_model.bin, clip/L14/pytorch_model.bin), while HUGGINGFACE_HUB_CACHE expects a models--<org>--<repo>/snapshots/<sha>/ tree. Leaving model.vision_encoder / model.clip_head at null sends asset resolution to hf_hub_download and fails under HF_HUB_OFFLINE=1 — point both at the files directly.
  • CI / skill-eval: stage weights + remapped JSON under $WORKSPACE_DIR from S3; do not rely on HuggingFace LFS downloads at eval time.

References

  • Shipped defaults: tao-pytorch nvidia_tao_pytorch/multimodal/video_clip/experiment_specs/ (when developing from source)
  • TAO Deploy workflow: references/tao-deploy-video-clip.md

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 16 other files (references) in skills/tao-finetune-video-clip of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • eval.config
  • evals/evals.json
  • references/skill_info.yaml
  • references/spec_template_deploy_evaluate.yaml
  • references/spec_template_deploy_inference.yaml
  • references/spec_template_evaluate.yaml
  • references/spec_template_export.yaml
  • references/spec_template_gen_trt_engine.yaml
  • references/spec_template_inference.yaml
  • references/spec_template_train.yaml
  • references/tao-deploy-video-clip.md
  • references/tao-deploy-video-clip.skill_info.yaml
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Finetune Video Clip next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Finetune Video Clip compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Finetune Video Clip this skillNVIDIA/skills3.6k—~3.5kAutomated safety check: NotesApache-2.0
Spark Environment Setupwshobson/agents40k—~2kAutomated safety check: PassMIT
Model Inference Optimizemajiayu000/spellbook287—~1.1kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
Onboard Jetpack5 Inference BackendsEGalahad/sim2real146—~1.1kAutomated safety check: PassNone
Llama CppOrchestra-Research/AI-Research-SKILLs13k3 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Model Inference Optimize

    majiayu000/spellbook

    优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

    287 GitHub stars~1.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

    146 GitHub stars~1.1k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Llama Cpp

    Orchestra-Research/AI-Research-SKILLs

    Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

    13k GitHub starsUsed in 3 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Builder

    qualcomm/qai-appbuilder

    QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder.

    247 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Finetune Video Clip

What does Tao Finetune Video Clip do?

InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment. Tao Finetune Video Clip is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment.

When should I use Tao Finetune Video Clip?

Tao Finetune Video Clip fits situations like: the user asks to fine-tune IV2CLIP; run videoclip train/evaluate/inference/export; build a Video-CLIP TensorRT engine; internVideo2-CLIP on KPI chunks.

How do I install Tao Finetune Video Clip in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-finetune-video-clip -a claude-code`. Or copy the skill folder (skills/tao-finetune-video-clip in NVIDIA/skills) into .claude/skills/tao-finetune-video-clip in your project. Claude Code loads it when a task matches its description.

How do I install Tao Finetune Video Clip in Codex?

Run `npx skills add NVIDIA/skills --skill tao-finetune-video-clip -a codex`. Or copy the skill folder (skills/tao-finetune-video-clip in NVIDIA/skills) into .agents/skills/tao-finetune-video-clip in your project. Codex loads it when a task matches its description.

Can I use Tao Finetune Video Clip in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-finetune-video-clip -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-finetune-video-clip, .gemini/skills/tao-finetune-video-clip, .github/skills/tao-finetune-video-clip and .opencode/skills/tao-finetune-video-clip in your project.

What does Tao Finetune Video Clip need to run?

Going by SKILL.md and its folder, Tao Finetune Video Clip needs the command-line tools its instructions call (docker and python) and credentials named NGC_KEY and HF_TOKEN. Our summary lists: Python 3; Docker; A credential in NGC_KEY. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit and the pinned TAO video_clip PyTorch and Deploy containers (see references/skill_info.yaml and references/tao-deploy-video-clip.skill_info.yaml), or a local tao-pytorch checkout + tao-cli venv for PyTorch virtualenv runs. MobileCLIP + InternVideo2 weights must be on disk for offline eval (HF LFS may be blocked in CI). Metadata JSON uses vadr1_chunks with absolute video_path entries..

Does Tao Finetune Video Clip access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tao Finetune Video Clip safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Finetune Video Clip use?

Tao Finetune Video Clip is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Finetune Video Clip use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4k tokens, read only when the agent opens those files.

What are the alternatives to Tao Finetune Video Clip?

Skills that share tags, products or a category with Tao Finetune Video Clip: Spark Environment Setup (wshobson/agents, 40k stars), Model Inference Optimize (majiayu000/spellbook, 287 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and Onboard Jetpack5 Inference Backends (EGalahad/sim2real, 146 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Finetune Video Clip?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.