Official agent skill

Tao Finetune Cosmos Embed

by NVIDIA in NVIDIA/skills

Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Tao Finetune Cosmos Embed

skills CLI
$ npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-finetune-cosmos-embed --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-finetune-cosmos-embed .claude/skills/tao-finetune-cosmos-embed && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-finetune-cosmos-embed
GitHub stars
3.6k
Token cost
~3.5k tokens
SKILL.md length
1,099 words
Files
12 (incl. references)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.

  • The user asks to fine-tune Cosmos-Embed1
  • SKILL.md covers Train Action Policy, Quick Start, Smoke Overrides and Data Format, plus 5 more sections
  • Calls docker, bash and python; needs HF_TOKEN
  • Run cosmos-embed inference

What it does

Tao Finetune Cosmos Embed is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning. Use when the user asks to "fine-tune Cosmos-Embed1", "run cosmos-embed inference", "export Cosmos-Embed1", "embed videos", or "search videos with text".

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading…

It sits in AI & LLM Engineering, covering Fine-tuning, Data cleaning and AI video generation. It works with Weights & Biases and gRPC. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user asks to fine-tune Cosmos-Embed1
  • Run cosmos-embed inference
  • Export Cosmos-Embed1
  • Search videos with text

Example prompts

  • “fine-tune Cosmos-Embed1”
  • “run cosmos-embed inference”
  • “export Cosmos-Embed1”
  • “/tao-finetune-cosmos-embed”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading pretrained `nvidia/Cosmos-Embed1-*` weights.
  • Pre-approved tools (allowed-tools): Read, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • bash
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading pretrained `nvidia/Cosmos-Embed1-*` weights.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Finetune Cosmos Embed loads about 3.5k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 1,099 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:74
    set -a; source /path/to/.env; set +a   # omit if already exported
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 1,099 words, ~3,452 tokens.

Download SKILL.mdSave it as .claude/skills/tao-finetune-cosmos-embed/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
tao-finetune-cosmos-embed
description
Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning. Use when the user asks to "fine-tune Cosmos-Embed1", "run cosmos-embed inference", "export Cosmos-Embed1", "embed videos", or "search videos with text".
allowed-tools
Read, Bash
compatibility
Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading pretrained `nvidia/Cosmos-Embed1-*` weights.
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
video, vision-language, vlm, multimodal, retrieval, embedding, cosmos, fine-tuning

Cosmos-Embed

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Cosmos-Embed1 is a joint video-text embedder for text-to-video retrieval, video-to-video search, zero-shot/kNN classification, and semantic deduplication. The packaged CLI is cosmos-embed1 and supports train, evaluate, inference, and export.

Container image and per-action commands are in references/skill_info.yaml. Compact starting specs are in references/spec_template_*.yaml.

Train Action Policy

AutoML is not packaged for this model skill because there are no Cosmos-Embed schemas under schemas/. Always use the direct model skill actions for train, evaluate, inference, and export, even when a higher-level request includes automl_policy: on. Do not route Cosmos-Embed through workflow or AutoML skills until model-specific train schemas and templates are added.

Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.

Quick Start

Use the published Cosmos-Embed container pinned below (also declared in references/skill_info.yaml). Do not build from the private Cosmos-Embed1 source tree for normal skill use; build from source only when developing the container itself.

bash
COSMOS_EMBED_IMAGE_DEFAULT=nvcr.io/nvidia/tao/tao-toolkit:7.1.0-cosmos-embed  # versions-key: images.tao_toolkit.cosmos_embed
COSMOS_EMBED_IMAGE="${COSMOS_EMBED_IMAGE:-$COSMOS_EMBED_IMAGE_DEFAULT}"
docker pull "$COSMOS_EMBED_IMAGE"

Expected local workspace layout:

text
workspace/
├── data/
│   ├── msrvtt_test_1k.json
│   └── video/
│       ├── video7020.mp4
│       └── ...
├── model/
│   └── Cosmos-Embed1-224p/        # optional if using HF repo id
├── specs/
│   ├── train.yaml
│   ├── evaluate.yaml
│   ├── inference.yaml
│   ├── export_onnx.yaml
│   └── export_hf.yaml
└── results/

Use these Docker options for all actions unless the local Docker/platform skill gives a stricter environment-specific command:

bash
set -a; source /path/to/.env; set +a   # omit if already exported
COSMOS_EMBED_IMAGE_DEFAULT=nvcr.io/nvidia/tao/tao-toolkit:7.1.0-cosmos-embed  # versions-key: images.tao_toolkit.cosmos_embed
COSMOS_EMBED_IMAGE="${COSMOS_EMBED_IMAGE:-$COSMOS_EMBED_IMAGE_DEFAULT}"
RUN_ROOT="${RUN_ROOT:-$PWD}"
DOCKER_COMMON=(
  --rm --gpus all --shm-size=8g --network=host
  --shm-size=64g
  --ulimit memlock=-1
  --ulimit stack=67108864
  -e HF_TOKEN
  -e WANDB_DISABLED=true
  -e WANDB_MODE=disabled
  -e HUGGINGFACE_HUB_CACHE=/hf_cache
  -v "$RUN_ROOT/data:/data:ro"
  -v "$RUN_ROOT/model:/model"
  -v "$RUN_ROOT/specs:/specs:ro"
  -v "$RUN_ROOT/results:/results"
  -v "$RUN_ROOT/hf_cache:/hf_cache"
)

For Cosmos-Embed images that ship protobuf==7.x, run a small startup preamble before every action:

bash
python -m pip install "protobuf<7"

The image contains wandb==0.21.0 with protobuf==7.x; importing W&B fails before training/evaluation unless protobuf is pinned below 7. Use WANDB_DISABLED=true and WANDB_MODE=disabled for smoke or offline runs. Cosmos-Embed may still download the public google-bert/bert-base-uncased Q-Former component even when the model checkpoint is disabled, so pass HF_TOKEN as an environment variable or mount a persistent HuggingFace cache. Do not write the token into specs, logs, or reports.

Train:

bash
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
  bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 train -e /specs/train.yaml results_dir=/results"

Evaluate:

bash
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
  bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 evaluate -e /specs/evaluate.yaml results_dir=/results"

Inference:

bash
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
  bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 inference -e /specs/inference.yaml \
  'inference.query.input_texts=[\"a man is singing on stage\"]' \
  inference.k=5 \
  results_dir=/results"

Export ONNX:

bash
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
  bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 export -e /specs/export_onnx.yaml \
  export.checkpoint=/results/train/checkpoints/iter_000000001.pt \
  export.onnx_file=/results/export/cosmos_embed1_combined.onnx \
  results_dir=/results"

Export HuggingFace format:

bash
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
  bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 export -e /specs/export_hf.yaml \
  export.checkpoint=/results/train/checkpoints/iter_000000001.pt \
  export.hf_output_dir=/results/export_hf/cosmos_embed1_hf \
  results_dir=/results"

Smoke Overrides

For a small functional check, keep the same specs and override the expensive knobs:

bash
train.max_iter=1
train.validation_iter=2
train.checkpoint_iter=1
train.optim.optim=adamw
train.optim.warmup_steps=0
train.optim.lr_decay_iters=1
dataset.train_dataset.batch_size=1
dataset.val_dataset.batch_size=1
dataset.train_dataset.workers=0
dataset.val_dataset.workers=0

When shortening the cosine scheduler for smoke runs, keep train.optim.lr_decay_iters greater than train.optim.warmup_steps, or set train.optim.warmup_steps=0 as shown above. The scheduler divides by lr_decay_iters - warmup_steps, so equal values fail before the checkpoint is written.

If no local Cosmos-Embed1 pretrained checkpoint is available, set model.pretrained_model_path=null for a plumbing-only smoke train. The model quality is meaningless in that mode, but the train/evaluate/inference/export action paths can still be exercised. In the current container, the Q-Former path can still fetch google-bert/bert-base-uncased; provide HF_TOKEN or a mounted HuggingFace cache for fresh ephemeral containers.

For evaluation and inference smoke tests on a tiny subset:

bash
evaluate.callbacks.embedding_visualization=false
evaluate.callbacks.max_eval_samples=8
dataset.test_dataset.batch_size=1
dataset.test_dataset.workers=0
inference.k=2
dataset.inference_dataset.batch_size=1
dataset.inference_dataset.workers=0

Data Format

The MSR-VTT path expects a local video glob and a JSON metadata file:

yaml
dataset:
  train_dataset:
    dataset_type: msrvtt
    mp4_urls: /data/video/*.mp4
    metadata: /data/msrvtt_test_1k.json

List-format metadata rows must include at least video and caption:

json
{"video_id": "video7020", "video": "video7020.mp4", "caption": "a woman creating a fondant baby and flower"}

The dataset loader derives the video id from the local .mp4 filename and filters to videos present in the metadata. If a run finds zero videos, check that mp4_urls points to a container-local glob and that metadata video names match the filenames.

Model Weights

  • Local HF directory: mount it under /model and set model.pretrained_model_path=/model/Cosmos-Embed1-224p.
  • HuggingFace repo: set model.pretrained_model_path=nvidia/Cosmos-Embed1-224p and pass HF_TOKEN if access is gated.
  • Fine-tuned checkpoint: set downstream actions to the resolver-selected /results/train/checkpoints/iter_#########.pt file.

Training writes full checkpoints under results/train/checkpoints/iter_#########.pt, updates results/train/checkpoints/latest_checkpoint.txt, and creates a cosmos_embed1_model_latest.pth symlink. For evaluate.checkpoint, inference.checkpoint, export.checkpoint, and train.resume_training_checkpoint_path, resolve and pass the exact iter_#########.pt file for the intended iteration. The action spec templates intentionally leave these checkpoint fields null so the model-skill runner or the user must provide the resolver-selected checkpoint. Use the latest symlink only when the user explicitly asks for latest.

For single-GPU resume/retrain from a consolidated checkpoint, set model.fsdp_shard_size: 1. The container default is 8, which sends resumed training through an FSDP apply path that Cosmos-Embed1 does not implement for this model class.

Variants:

VariantResolutionFramesEmbedding dim
Cosmos-Embed1-224p224 x 2248256
Cosmos-Embed1-336p336 x 3368768
Cosmos-Embed1-448p448 x 4488768

Keep model.network.embed_dim, model.input_hw, and model.network.spatial_resolution aligned with the selected variant.

Show full SKILL.md (463 more words)Show less

Important Parameters

ParameterNotes
train.num_gpus1 for single GPU, >1 auto-launches torchrun, -1 auto-detects visible GPUs.
train.max_iterMain training length. Use 1 only for smoke testing.
train.optim.optimfused_adamw is faster when available; adamw is safer for smoke and portability.
model.lora.enabledEnables LoRA. Set model.network.visual_encoder.transformer_engine=false when LoRA is on.
model.lora.lora_rankLoRA rank. Start with 8; try 4, 8, or 16 for manual or AutoML-style sweeps.
model.lora.lora_alphaLoRA scaling factor. Start with 16; keep near 2 * lora_rank unless experiments show otherwise.
model.lora.lora_dropoutLoRA dropout. Start with 0.1; sweep 0.0, 0.05, and 0.1 for small datasets.
model.lora.biasBias policy: none, all, or lora_only. Keep none unless intentionally training biases.
model.lora.use_rslora / use_doraOptional LoRA variants. Enable one at a time and record the setting with the checkpoint.
model.lora.target_modulesOptional module-name patterns for LoRA injection. Leave empty for the default ViT + Q-Former attention/MLP targets.
model.lora.modules_to_saveOptional modules to keep fully trainable alongside LoRA. Leave empty unless preserving a task-specific head.
evaluate.load_dataset_pkl / save_dataset_pklCache evaluation embeddings.
inference.load_dataset_pkl / save_dataset_pklCache the search database for repeated retrieval.
export.modevideo, text, combined, or huggingface.
export.on_cpuRecommended for export to avoid device mismatch issues.
LoRA and AutoML Notes

For parameter-efficient fine-tuning, set model.lora.enabled=true and keep model.network.visual_encoder.transformer_engine=false; TAO Core's Cosmos-Embed1 config notes that PEFT cannot inject adapters into Transformer Engine layers. Treat the LoRA fields above as the first candidate parameters for manual tuning or AutoML-style search before unfreezing larger model blocks. Avoid changing target_modules or modules_to_save unless the user explicitly needs custom adapter placement.

S3 Staging

The Cosmos-Embed1 CLI consumes local paths and Python globs, not raw s3://.../*.mp4 URIs. For S3-backed runs, first stage a subset or full dataset to the execution host/container filesystem, then use local paths such as /data/video/*.mp4 in the spec.

Recommended S3 layout for staged MSR-VTT data:

text
s3://bucket/path/cosmos-embed/msrvtt-subset/
├── msrvtt_test_1k.json
└── video/
    ├── video7020.mp4
    └── ...

After downloading/syncing that prefix into the mounted data/ directory, use the same Docker commands above.

Outputs

text
results/
├── train/
│   ├── cosmos_embed1_model_latest.pth
│   ├── cosmos_embed1_model_<iter>.pth
│   └── experiment.yaml
├── evaluate/
│   ├── metrics.json
│   └── experiment.yaml
├── inference/
│   ├── results.json
│   └── experiment.yaml
├── export/
│   ├── cosmos_embed1_combined.onnx
│   └── export_config.yaml
└── export_hf/
    └── cosmos_embed1_hf/

Known Pitfalls

SymptomCauseFix
MSRVTTDataset: 0 videos foundmp4_urls is not a local glob or metadata filenames do not match videos.Mount data into the container and set mp4_urls=/data/video/*.mp4.
HF download/auth failureMissing or invalid HF_TOKEN, or model agreement not accepted.Accept the model terms and pass -e HF_TOKEN.
cannot import name 'Imports' from 'wandb.proto.wandb_telemetry_pb2'wandb==0.21.0 in the container is incompatible with protobuf==7.x.Run python -m pip install "protobuf<7" in the container before invoking cosmos-embed1.
Resume fails with Model does not implement 'apply_fsdp'Single-GPU resume loaded a consolidated checkpoint while model.fsdp_shard_size stayed at the default 8.Set model.fsdp_shard_size=1 for local single-GPU resume/retrain.
LoRA injection failureTransformer Engine visual encoder is enabled.Set model.network.visual_encoder.transformer_engine=false.
ONNX/HF export complains about missing componentsExport checkpoint is partial or adapter-only.Use a full checkpoint or configure pretrained visual/text sources before export.
CUDA OOMBatch/resolution too high for the GPU.Reduce batch size, use 224p, enable LoRA, or use more GPUs.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in skills/tao-finetune-cosmos-embed of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/skill_info.yaml
  • references/spec_template_evaluate.yaml
  • references/spec_template_export_hf.yaml
  • references/spec_template_export_onnx.yaml
  • references/spec_template_inference.yaml
  • references/spec_template_train.yaml
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Finetune Cosmos Embed next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Finetune Cosmos Embed compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Finetune Cosmos Embed this skillNVIDIA/skills3.6k—~3.5kAutomated safety check: NotesApache-2.0
Gemini Live APIgoogle/skills21k—~2.5kAutomated safety check: NotesApache-2.0
Swift Mlx Lmkellyvv/PhoneClaw1.3k—~3.7kAutomated safety check: PassApache-2.0
Civitaiartokun/comfyui-mcp803—~1.1kAutomated safety check: PassMIT
Wan T2v Videoartokun/comfyui-mcp803—~3.3kAutomated safety check: PassMIT
Sft Launchopen-thoughts/OpenThoughts-Agent301—~2.9kAutomated safety check: PassApache-2.0

Similar skills

  • Gemini Live API

    google/skills

    Official

    Generates a Gemini LiveAPI client service class in the user's chosen programming language.

    21k GitHub stars~2.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Swift Mlx Lm

    kellyvv/PhoneClaw

    MLX Swift LM - Run LLMs and VLMs on Apple Silicon using MLX.

    1.3k GitHub stars~3.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Civitai

    artokun/comfyui-mcp

    Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.

    803 GitHub stars~1.1k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Wan T2v Video

    artokun/comfyui-mcp

    Build WAN 2.2 Text-to-Video workflows. An agent skill from artokun/comfyui-mcp.

    803 GitHub stars~3.3k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Sft Launch

    open-thoughts/OpenThoughts-Agent

    Launch SFT via python -m hpc.launch --jobtype sft on any cluster (JSC Jupiter GH200, CINECA Leonardo A100, TACC Vista GH200), with EITHER backend — LLaMA-Factory (default) or axolotl (--sftbackend…

    301 GitHub stars~2.9k tokensUpdated 12 days ago
    AI & LLM EngineeringAuto-check passed
  • Discover ML

    rand/cc-polymath

    Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.

    181 GitHub stars~574 tokensUpdated 7 mo ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Finetune Cosmos Embed

What does Tao Finetune Cosmos Embed do?

Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning. Tao Finetune Cosmos Embed is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.

When should I use Tao Finetune Cosmos Embed?

Tao Finetune Cosmos Embed fits situations like: the user asks to fine-tune Cosmos-Embed1; run cosmos-embed inference; export Cosmos-Embed1; search videos with text.

How do I install Tao Finetune Cosmos Embed in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a claude-code`. Or copy the skill folder (skills/tao-finetune-cosmos-embed in NVIDIA/skills) into .claude/skills/tao-finetune-cosmos-embed in your project. Claude Code loads it when a task matches its description.

How do I install Tao Finetune Cosmos Embed in Codex?

Run `npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a codex`. Or copy the skill folder (skills/tao-finetune-cosmos-embed in NVIDIA/skills) into .agents/skills/tao-finetune-cosmos-embed in your project. Codex loads it when a task matches its description.

Can I use Tao Finetune Cosmos Embed in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-finetune-cosmos-embed, .gemini/skills/tao-finetune-cosmos-embed, .github/skills/tao-finetune-cosmos-embed and .opencode/skills/tao-finetune-cosmos-embed in your project.

What does Tao Finetune Cosmos Embed need to run?

Going by SKILL.md and its folder, Tao Finetune Cosmos Embed needs the command-line tools its instructions call (docker, bash and python) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading pretrained `nvidia/Cosmos-Embed1-*` weights..

Does Tao Finetune Cosmos Embed access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tao Finetune Cosmos Embed safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Finetune Cosmos Embed use?

Tao Finetune Cosmos Embed is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Finetune Cosmos Embed use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Tao Finetune Cosmos Embed?

Skills that share tags, products or a category with Tao Finetune Cosmos Embed: Gemini Live API (google/skills, 21k stars), Swift Mlx Lm (kellyvv/PhoneClaw, 1.3k stars), Civitai (artokun/comfyui-mcp, 803 stars) and Wan T2v Video (artokun/comfyui-mcp, 803 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Finetune Cosmos Embed?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.