Official agent skill

Tao Finetune Clip

by NVIDIA in NVIDIA/skills

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Finetune Clip

skills CLI
$ npx skills add NVIDIA/skills --skill tao-finetune-clip -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-finetune-clip --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-finetune-clip .claude/skills/tao-finetune-clip && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-finetune-clip
GitHub stars
3.6k
Token cost
~4k tokens
SKILL.md length
1,541 words
Files
20 (incl. references)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

  • Running zero-shot classification
  • SKILL.md covers Train Action Policy, Instructions, Training Requirements and Eval Dataset, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Computing image embeddings

What it does

Tao Finetune Clip is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment. Use when fine-tuning or training CLIP, running zero-shot classification, computing image embeddings, or deploying CLIP to ONNX/TensorRT. This is a single-action model skill; do not use it for an iterative weak-attribute improvement loop that keeps retraining and evaluating until progress stops, which belongs to tao-run-deft-pas.

Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 23 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.

It sits in AI & LLM Engineering, covering Fine-tuning, Embeddings and LLM inference and serving. It works with NVIDIA AI Platform, ONNX and PyTorch. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Running zero-shot classification
  • Computing image embeddings
  • Deploying CLIP to ONNX/TensorRT
  • An iterative weak-attribute improvement loop that keeps retraining and evaluating until progress stops

Example prompts

  • “/tao-finetune-clip”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit.
  • Pre-approved tools (allowed-tools): Read, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Finetune Clip loads about 4k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 122 tokens; SKILL.md has 1,541 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 1,541 words, ~3,968 tokens.

Download SKILL.mdSave it as .claude/skills/tao-finetune-clip/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.
name
tao-finetune-clip
description
CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment. Use when fine-tuning or training CLIP, running zero-shot classification, computing image embeddings, or deploying CLIP to ONNX/TensorRT. This is a single-action model skill; do not use it for an iterative weak-attribute improvement loop that keeps retraining and evaluating until progress stops, which belongs to tao-run-deft-pas.
allowed-tools
Read, Bash
compatibility
Requires docker + nvidia-container-toolkit.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
vision-language, classification, embedding, zero-shot, deployment

CLIP

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Contrastive Language-Image Pre-training model for zero-shot and fine-tuned image classification, image-text retrieval, and embedding extraction. Fine-tuning adapts CLIP's shared image-text embedding space to domain-specific image-caption data.

No default NGC pretrained checkpoint is required for spec construction, but unset checkpoint behavior is action-specific. In the validation-fixes PyTorch image, export.checkpoint: null exports the selected CLIP architecture and may initialize weights when pretrained weights are unavailable. Do not assume inference.checkpoint: null loads pretrained weights: clip inference currently calls the checkpoint loader with None and fails before embedding extraction. For PyTorch inference, checkpoint-backed evaluation/export, resume, and retrain flows, resolve and pass an exact checkpoint from the parent train output. For trusted TAO checkpoints produced by the current run or a known parent job, set TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1 on checkpoint-dependent PyTorch actions so PyTorch 2.6 can load the Lightning checkpoint metadata; do not set this for untrusted checkpoints.

Supported actions: train, evaluate, inference, export, gen_trt_engine.

Train Action Policy

This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.

The packaged CLIP train schema enables train.optim.vision_lr and train.optim.text_lr as default AutoML search parameters. For smoke tests, keep the search small by using the Bayesian algorithm with two recommendations and narrow LR ranges.

Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.

Instructions

Use this skill for NVIDIA TAO CLIP jobs: training, evaluation, embedding inference, ONNX export, and TensorRT engine generation. Start by identifying the requested action, then load only the referenced files needed for that action: defaults.json for default parameters, config.json for action/data-source wiring, references/spec_template.yaml for full spec shape, and references/model_info.yaml for SDK metadata.

For dataset-backed actions, collect the required image, caption, list, or prompt files from the user and place the resolved paths in spec_overrides. For local Docker runs, mount extracted folders in the container and point image_dir / caption_dir at those folders; if a data source provides .tar.gz archives, extract them before running the in-container CLIP commands. For export and gen_trt_engine, infer parent artifacts from the upstream job when available; otherwise require explicit checkpoint, ONNX, or engine paths. Run gen_trt_engine, TensorRT evaluate, and TensorRT inference in the TAO Deploy image.

For TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and TensorRT inference), read references/tao-deploy-clip.md first. Deploy spec templates live in this skill's references/ folder with the spec_template_deploy_*.yaml prefix.

Training Requirements

  • Dataset type: image_text
  • Formats: custom image/caption folders or WebDataset shards
  • Monitoring metric: val/t2i_mAP

The train action emits val/t2i_mAP, which is the AutoML selection objective. The standalone evaluate action reports the corresponding held-out metric as test/t2i_mAP; use that name for checkpoint evaluation and compare its value with the selected training validation metric rather than expecting a val/ key from the evaluate action.

Supported Models
  • OpenCLIP / NV-CLIP: ViT-L-14-SigLIP-CLIPA-224 (default), ViT-L-14-SigLIP-CLIPA-336, ViT-H-14-SigLIP-CLIPA-224, ViT-H-14-SigLIP-CLIPA-336, ViT-H-14-SigLIP-CLIPA-574
  • Radio-CLIP: c-radio_v3-b, c-radio_v3-l, c-radio_v3-h, c-radio_v3-g
  • SigLIP2: siglip2-so400m-patch16-256, siglip2-so400m-patch14-224, siglip2-so400m-patch14-384, siglip2-so400m-patch16-384, siglip2-so400m-patch16-512, siglip2-so400m-patch16-naflex

Radio-CLIP requires model.adaptor_name to be set to siglip or clip.

Per-Action Dataset Requirements
ActionSpec KeySourceFilesList?
traindataset.train.datasetstrain_datasetsimage_dir: images.tar.gz, image_list_file: image_list.txt, caption_dir: captions.tar.gzYes
traindataset.train.wds.root_dirtrain_wds_datasetroot directory containing .tar shardsNo
traindataset.train.wds.shard_list_filetrain_wds_datasetshards.txt listing shard pathsNo
traindataset.val.datasetseval_datasetimage_dir: images.tar.gz, image_list_file: image_list.txt, caption_dir: captions.tar.gzYes
evaluatedataset.val.datasetseval_datasetimage_dir: images.tar.gz, image_list_file: image_list.txt, caption_dir: captions.tar.gzYes
inferenceinference.datasetsinference_datasetimage_dir: images.tar.gzYes
inferenceinference.text_fileinference_datasetprompts.txtNo
exportexport.checkpointparent train job or explicit checkpointcheckpoint .pth, optional for pretrained exportNo
gen_trt_enginegen_trt_engine.onnx_fileparent export job or explicit ONNXclip_model.onnxNo

For custom training, set dataset.train.type: custom and provide dataset.train.datasets entries. Image and caption files must share the same base name. caption_file_suffix defaults to .txt, and image_list_file is optional.

When no native CLIP image-caption dataset is available, do not silently treat image-classification data as CLIP data. If the user explicitly allows a plumbing-only validation fallback, derive caption files from class labels, document that the captions are generated from labels, and keep each image/caption pair on the same base filename. Without an image_list_file, the TAO custom loader scans the configured image directory for image files; keep validation folders flat unless you provide a list file.

For WDS training, set dataset.train.type: wds and provide at least one of dataset.train.wds.root_dir or dataset.train.wds.shard_list_file. root_dir is scanned recursively for .tar shards. shard_list_file is a text file with one shard path per line; relative lines resolve under the list-file directory unless root_dir is also supplied, in which case they resolve under root_dir. Validation/evaluation data remains custom format via dataset.val.datasets.

Show full SKILL.md (655 more words)Show less
Typical Spec Overrides

Data source overrides are mandatory for dataset-backed actions. Construct paths from the Per-Action Dataset Requirements table and include them in spec_overrides. For inference, provide at least one of inference.datasets or inference.text_file.

python
S3_TRAIN = "s3://bucket/data/train"
S3_WDS = "s3://bucket/data/wds"
S3_EVAL = "s3://bucket/data/eval"
S3_INFER = "s3://bucket/data/infer"

train, custom dataset:

python
{
    "train.num_epochs": 10,
    "dataset.train.type": "custom",
    "dataset.train.datasets": [{"image_dir": f"{S3_TRAIN}/images.tar.gz", "image_list_file": f"{S3_TRAIN}/image_list.txt", "caption_dir": f"{S3_TRAIN}/captions.tar.gz"}],
    "dataset.val.datasets": [{"image_dir": f"{S3_EVAL}/images.tar.gz", "image_list_file": f"{S3_EVAL}/image_list.txt", "caption_dir": f"{S3_EVAL}/captions.tar.gz"}],
}

train, WDS dataset:

python
{
    "train.num_epochs": 10,
    "dataset.train.type": "wds",
    "dataset.train.wds.root_dir": f"{S3_WDS}",
    "dataset.train.wds.shard_list_file": f"{S3_WDS}/shards.txt",
    "dataset.train.wds.samples_per_shard": 10000,
    "dataset.val.datasets": [{"image_dir": f"{S3_EVAL}/images.tar.gz", "image_list_file": f"{S3_EVAL}/image_list.txt", "caption_dir": f"{S3_EVAL}/captions.tar.gz"}],
}

evaluate:

python
{
    "dataset.val.datasets": [{"image_dir": f"{S3_EVAL}/images.tar.gz", "image_list_file": f"{S3_EVAL}/image_list.txt", "caption_dir": f"{S3_EVAL}/captions.tar.gz"}],
}

Leave evaluate.checkpoint unset for zero-shot evaluation with pretrained weights. Set evaluate.trt_engine instead of evaluate.checkpoint for TensorRT evaluation.

inference:

python
{
    "inference.datasets": [{"image_dir": f"{S3_INFER}/images.tar.gz"}],
    "inference.text_file": f"{S3_INFER}/prompts.txt",
}

Inference writes image_embeddings.h5 and/or text_embeddings.h5 under results_dir. The saved embeddings are L2-normalized.

export:

python
{
    "export.onnx_file": "${results_dir}/export/clip_model.onnx",
    "export.encoder_type": "combined",
    "export.batch_size": -1,
}

Set export.encoder_type: separate when deployment should use independent vision and text encoders. Separate export writes _vision.onnx and _text.onnx variants derived from the base export.onnx_file.

For checkpoint-dependent actions, use the model-specific checkpoint resolver output from the parent train job. CLIP training writes checkpoints such as model_epoch_000_step_00020.pth and a clip_latest.pth symlink. Use the exact resolved checkpoint for evaluate.checkpoint, inference.checkpoint, export.checkpoint, and train.resume_training_checkpoint_path; use clip_latest.pth only when the user explicitly asks for latest.

When the resolved checkpoint is trusted TAO output, checkpoint-backed PyTorch evaluate, inference, export, and resume training should run with TORCH_FORCE_NO_WEIGHTS_ONLY_LOAD=1. PyTorch 2.6 otherwise defaults checkpoint loading to weights-only mode and can reject CLIP Lightning checkpoints containing NumPy scalar metadata.

gen_trt_engine:

python
{
    "gen_trt_engine.onnx_file": "${results_dir}/export/clip_model.onnx",
    "gen_trt_engine.trt_engine": "${results_dir}/deploy/clip_model.engine",
    "gen_trt_engine.batch_size": -1,
    "gen_trt_engine.tensorrt.data_type": "fp16",
    "gen_trt_engine.tensorrt.min_batch_size": 1,
    "gen_trt_engine.tensorrt.opt_batch_size": 1,
    "gen_trt_engine.tensorrt.max_batch_size": 16,
}

Eval Dataset

Optional for training. If provided, validation metrics are computed at validation intervals. Required for evaluate.

Deploy Workflow

The skill exposes gen_trt_engine as the deploy action. In generated SDK runners, use model_info["actions"]["gen_trt_engine"] and run it in the TAO Deploy image, not the PyTorch training image. The in-container command is clip gen_trt_engine -e {config_path}; direct TAO Launcher usage spells the same action as tao deploy clip gen_trt_engine -e /path/to/spec.yaml.

TAO Deploy inference can discover combined engines, paired separate engines, or single-pillar _vision.engine / _text.engine files. For full TensorRT retrieval evaluation or image+text TensorRT inference, export with export.encoder_type: separate and run clip gen_trt_engine twice: build clip_model_vision.onnx to an engine ending in _vision.engine, then build clip_model_text.onnx to the matching _text.engine in the same directory. For image-only TensorRT inference, building only the _vision.engine is sufficient and inference.text_file must be null. TensorRT evaluate and text inference require a text-capable engine; if only a vision engine is present, deploy evaluation fails because text embeddings cannot be extracted.

Use evaluate.trt_engine for TensorRT evaluation and inference.trt_engine for TensorRT embedding extraction. These TensorRT paths also run in the TAO Deploy image. Direct TAO Launcher usage spells these as tao deploy clip evaluate and tao deploy clip inference.

Important Parameters

  • model.type: Backbone family and resolution. Use a TAO-registered CLIP model ID such as ViT-L-14-SigLIP-CLIPA-224. Prefer the listed OpenCLIP / NV-CLIP IDs for AutoML smoke tests because the current TAO container registry routes them through the supported augmentation adapter.
  • model.adaptor_name: Required for Radio-CLIP. Set to siglip or clip.
  • model.image_size: Training transform image resolution. Keep it aligned with the selected fixed-resolution backbone.
  • train.num_epochs: CLIP fine-tuning often converges quickly. Start with 10-20 epochs for domain adaptation, then increase only if validation loss is still improving.
  • train.optim.vision_lr / train.optim.text_lr: Learning rates for the two encoders. CLIP is sensitive to high learning rates; reduce both if loss is unstable.
  • model.freeze_vision_encoder / model.freeze_text_encoder: Defaults are false. Freezing one encoder can help when the dataset is small or only one modality needs adaptation.
  • train.loss_type: siglip is recommended for SigLIP2 and Radio-CLIP. Use clip for CLIP-style softmax loss.
  • export.encoder_type: combined exports one ONNX graph. separate exports independent vision and text graphs.
  • gen_trt_engine.tensorrt.data_type: TensorRT deployment supports fp16 and fp32.

Hardware

Single-GPU training works for small datasets. Use 4+ GPUs for datasets with more than 100k images or large backbones. Use 16GB+ VRAM per GPU for small/fixed-resolution runs and larger GPUs for Radio-CLIP or high-resolution OpenCLIP variants.

Error Patterns

See references/error-patterns.md for the full list of CLIP error symptoms and fixes (CUDA OOM, NaN loss, retrieval quality, dataset format/size, Radio-CLIP and model-ID validation, ONNX external data, TensorRT shape mismatch, PyTorch 2.6 checkpoint load, null-checkpoint inference, TensorRT text/retrieval failures, attention_mask handling, and spec/schema merge errors).

Spec Param / Parent Model Inference

See references/spec-param-inference.md for the model-specific inference mappings (the full clip.config.json action/spec-field/inference-function table) that generated runners apply with SDK helpers before create_job(), plus parent_job_id resolution rules.

Deployment

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 19 other files (references) in skills/tao-finetune-clip of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/error-patterns.md
  • references/skill_info.yaml
  • references/spec-param-inference.md
  • references/spec_template.yaml
  • references/spec_template_deploy.yaml
  • references/spec_template_evaluate.yaml
  • references/spec_template_export.yaml
  • references/spec_template_train.yaml
  • references/tao-deploy-clip.md
  • references/tao-deploy-clip.skill_info.yaml
  • schemas/evaluate.schema.json
  • schemas/export.schema.json
  • schemas/manifest.json
  • … and 3 more

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Finetune Clip next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Finetune Clip compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Finetune Clip this skillNVIDIA/skills3.6k—~4kAutomated safety check: NotesApache-2.0
Senior Computer Visionalirezarezvani/claude-skills28k1 repos~3.2kAutomated safety check: PassMIT
Spark Environment Setupwshobson/agents40k—~2kAutomated safety check: PassMIT
Model Inference Optimizemajiayu000/spellbook287—~1.1kAutomated safety check: PassMIT
Matlab Use Visual Inspectionmatlab/matlab-agentic-toolkit1.1k—~3.1kAutomated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0

Similar skills

  • Senior Computer Vision

    alirezarezvani/claude-skills

    Computer vision engineering skill for object detection, image segmentation, and visual AI systems.

    28k GitHub starsUsed in 1 repo~3.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Model Inference Optimize

    majiayu000/spellbook

    优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

    287 GitHub stars~1.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Matlab Use Visual Inspection

    matlab/matlab-agentic-toolkit

    Build machine vision inspection systems with MATLAB Visual Inspection Toolbox.

    1.1k GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable.

    146 GitHub stars~1.1k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Finetune Clip

What does Tao Finetune Clip do?

CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment. Tao Finetune Clip is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment.

When should I use Tao Finetune Clip?

Tao Finetune Clip fits situations like: running zero-shot classification; computing image embeddings; deploying CLIP to ONNX/TensorRT; an iterative weak-attribute improvement loop that keeps retraining and evaluating until progress stops.

How do I install Tao Finetune Clip in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-finetune-clip -a claude-code`. Or copy the skill folder (skills/tao-finetune-clip in NVIDIA/skills) into .claude/skills/tao-finetune-clip in your project. Claude Code loads it when a task matches its description.

How do I install Tao Finetune Clip in Codex?

Run `npx skills add NVIDIA/skills --skill tao-finetune-clip -a codex`. Or copy the skill folder (skills/tao-finetune-clip in NVIDIA/skills) into .agents/skills/tao-finetune-clip in your project. Codex loads it when a task matches its description.

Can I use Tao Finetune Clip in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-finetune-clip -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-finetune-clip, .gemini/skills/tao-finetune-clip, .github/skills/tao-finetune-clip and .opencode/skills/tao-finetune-clip in your project.

What does Tao Finetune Clip need to run?

SKILL.md names no scripts, command-line tools or credentials: Tao Finetune Clip is instructions for the agent only. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..

Does Tao Finetune Clip access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tao Finetune Clip safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Finetune Clip use?

Tao Finetune Clip is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Finetune Clip use?

About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.6k tokens, read only when the agent opens those files.

What are the alternatives to Tao Finetune Clip?

Skills that share tags, products or a category with Tao Finetune Clip: Senior Computer Vision (alirezarezvani/claude-skills, 28k stars), Spark Environment Setup (wshobson/agents, 40k stars), Model Inference Optimize (majiayu000/spellbook, 287 stars) and Matlab Use Visual Inspection (matlab/matlab-agentic-toolkit, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Finetune Clip?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.