Gemini Live API
google/skills
Generates a Gemini LiveAPI client service class in the user's chosen programming language.
Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.
$ npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-finetune-cosmos-embed --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-finetune-cosmos-embed .claude/skills/tao-finetune-cosmos-embed && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-finetune-cosmos-embed" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-cosmos-embed into .claude/skills/tao-finetune-cosmos-embed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-finetune-cosmos-embed", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-cosmos-embedType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-finetune-cosmos-embed --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-finetune-cosmos-embed .agents/skills/tao-finetune-cosmos-embed && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-finetune-cosmos-embed" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-cosmos-embed into .agents/skills/tao-finetune-cosmos-embed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-finetune-cosmos-embed", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-finetune-cosmos-embed --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-finetune-cosmos-embed .cursor/skills/tao-finetune-cosmos-embed && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-finetune-cosmos-embed" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-cosmos-embed into .cursor/skills/tao-finetune-cosmos-embed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-finetune-cosmos-embed", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-finetune-cosmos-embed--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-finetune-cosmos-embed --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-finetune-cosmos-embed .gemini/skills/tao-finetune-cosmos-embed && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-finetune-cosmos-embed" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-cosmos-embed into .gemini/skills/tao-finetune-cosmos-embed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-finetune-cosmos-embed", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-finetune-cosmos-embedInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-finetune-cosmos-embed .github/skills/tao-finetune-cosmos-embed && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-finetune-cosmos-embed" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-cosmos-embed into .github/skills/tao-finetune-cosmos-embed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-finetune-cosmos-embed", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-finetune-cosmos-embed --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-finetune-cosmos-embed .opencode/skills/tao-finetune-cosmos-embed && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-finetune-cosmos-embed" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-finetune-cosmos-embed into .opencode/skills/tao-finetune-cosmos-embed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-finetune-cosmos-embed", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-finetune-cosmos-embedCosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.
Tao Finetune Cosmos Embed is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning. Use when the user asks to "fine-tune Cosmos-Embed1", "run cosmos-embed inference", "export Cosmos-Embed1", "embed videos", or "search videos with text".
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading…
It sits in AI & LLM Engineering, covering Fine-tuning, Data cleaning and AI video generation. It works with Weights & Biases and gRPC. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockerbashpythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading pretrained `nvidia/Cosmos-Embed1-*` weights.
From compatibility in the SKILL.md frontmatter.
Tao Finetune Cosmos Embed loads about 3.5k tokens when it runs, and up to ~5.7k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 1,099 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
set -a; source /path/to/.env; set +a # omit if already exportedallowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 1,099 words, ~3,452 tokens.
.claude/skills/tao-finetune-cosmos-embed/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Cosmos-Embed1 is a joint video-text embedder for text-to-video retrieval, video-to-video search, zero-shot/kNN classification, and semantic deduplication. The packaged CLI is cosmos-embed1 and supports train, evaluate, inference, and export.
Container image and per-action commands are in references/skill_info.yaml. Compact starting specs are in references/spec_template_*.yaml.
AutoML is not packaged for this model skill because there are no Cosmos-Embed schemas under schemas/. Always use the direct model skill actions for train, evaluate, inference, and export, even when a higher-level request includes automl_policy: on. Do not route Cosmos-Embed through workflow or AutoML skills until model-specific train schemas and templates are added.
Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.
Use the published Cosmos-Embed container pinned below (also declared in
references/skill_info.yaml). Do not build from the private
Cosmos-Embed1 source tree for normal skill use; build from source only when
developing the container itself.
COSMOS_EMBED_IMAGE_DEFAULT=nvcr.io/nvidia/tao/tao-toolkit:7.1.0-cosmos-embed # versions-key: images.tao_toolkit.cosmos_embed
COSMOS_EMBED_IMAGE="${COSMOS_EMBED_IMAGE:-$COSMOS_EMBED_IMAGE_DEFAULT}"
docker pull "$COSMOS_EMBED_IMAGE"Expected local workspace layout:
workspace/
├── data/
│ ├── msrvtt_test_1k.json
│ └── video/
│ ├── video7020.mp4
│ └── ...
├── model/
│ └── Cosmos-Embed1-224p/ # optional if using HF repo id
├── specs/
│ ├── train.yaml
│ ├── evaluate.yaml
│ ├── inference.yaml
│ ├── export_onnx.yaml
│ └── export_hf.yaml
└── results/Use these Docker options for all actions unless the local Docker/platform skill gives a stricter environment-specific command:
set -a; source /path/to/.env; set +a # omit if already exported
COSMOS_EMBED_IMAGE_DEFAULT=nvcr.io/nvidia/tao/tao-toolkit:7.1.0-cosmos-embed # versions-key: images.tao_toolkit.cosmos_embed
COSMOS_EMBED_IMAGE="${COSMOS_EMBED_IMAGE:-$COSMOS_EMBED_IMAGE_DEFAULT}"
RUN_ROOT="${RUN_ROOT:-$PWD}"
DOCKER_COMMON=(
--rm --gpus all --shm-size=8g --network=host
--shm-size=64g
--ulimit memlock=-1
--ulimit stack=67108864
-e HF_TOKEN
-e WANDB_DISABLED=true
-e WANDB_MODE=disabled
-e HUGGINGFACE_HUB_CACHE=/hf_cache
-v "$RUN_ROOT/data:/data:ro"
-v "$RUN_ROOT/model:/model"
-v "$RUN_ROOT/specs:/specs:ro"
-v "$RUN_ROOT/results:/results"
-v "$RUN_ROOT/hf_cache:/hf_cache"
)For Cosmos-Embed images that ship protobuf==7.x, run a small startup
preamble before every action:
python -m pip install "protobuf<7"The image contains wandb==0.21.0 with protobuf==7.x; importing W&B fails before training/evaluation unless protobuf is pinned below 7. Use WANDB_DISABLED=true and WANDB_MODE=disabled for smoke or offline runs. Cosmos-Embed may still download the public google-bert/bert-base-uncased Q-Former component even when the model checkpoint is disabled, so pass HF_TOKEN as an environment variable or mount a persistent HuggingFace cache. Do not write the token into specs, logs, or reports.
Train:
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 train -e /specs/train.yaml results_dir=/results"Evaluate:
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 evaluate -e /specs/evaluate.yaml results_dir=/results"Inference:
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 inference -e /specs/inference.yaml \
'inference.query.input_texts=[\"a man is singing on stage\"]' \
inference.k=5 \
results_dir=/results"Export ONNX:
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 export -e /specs/export_onnx.yaml \
export.checkpoint=/results/train/checkpoints/iter_000000001.pt \
export.onnx_file=/results/export/cosmos_embed1_combined.onnx \
results_dir=/results"Export HuggingFace format:
docker run "${DOCKER_COMMON[@]}" "$COSMOS_EMBED_IMAGE" \
bash -lc "python -m pip install 'protobuf<7' && cosmos-embed1 export -e /specs/export_hf.yaml \
export.checkpoint=/results/train/checkpoints/iter_000000001.pt \
export.hf_output_dir=/results/export_hf/cosmos_embed1_hf \
results_dir=/results"For a small functional check, keep the same specs and override the expensive knobs:
train.max_iter=1
train.validation_iter=2
train.checkpoint_iter=1
train.optim.optim=adamw
train.optim.warmup_steps=0
train.optim.lr_decay_iters=1
dataset.train_dataset.batch_size=1
dataset.val_dataset.batch_size=1
dataset.train_dataset.workers=0
dataset.val_dataset.workers=0When shortening the cosine scheduler for smoke runs, keep
train.optim.lr_decay_iters greater than train.optim.warmup_steps, or set
train.optim.warmup_steps=0 as shown above. The scheduler divides by
lr_decay_iters - warmup_steps, so equal values fail before the checkpoint is
written.
If no local Cosmos-Embed1 pretrained checkpoint is available, set model.pretrained_model_path=null for a plumbing-only smoke train. The model quality is meaningless in that mode, but the train/evaluate/inference/export action paths can still be exercised. In the current container, the Q-Former path can still fetch google-bert/bert-base-uncased; provide HF_TOKEN or a mounted HuggingFace cache for fresh ephemeral containers.
For evaluation and inference smoke tests on a tiny subset:
evaluate.callbacks.embedding_visualization=false
evaluate.callbacks.max_eval_samples=8
dataset.test_dataset.batch_size=1
dataset.test_dataset.workers=0
inference.k=2
dataset.inference_dataset.batch_size=1
dataset.inference_dataset.workers=0The MSR-VTT path expects a local video glob and a JSON metadata file:
dataset:
train_dataset:
dataset_type: msrvtt
mp4_urls: /data/video/*.mp4
metadata: /data/msrvtt_test_1k.jsonList-format metadata rows must include at least video and caption:
{"video_id": "video7020", "video": "video7020.mp4", "caption": "a woman creating a fondant baby and flower"}The dataset loader derives the video id from the local .mp4 filename and filters to videos present in the metadata. If a run finds zero videos, check that mp4_urls points to a container-local glob and that metadata video names match the filenames.
/model and set model.pretrained_model_path=/model/Cosmos-Embed1-224p.model.pretrained_model_path=nvidia/Cosmos-Embed1-224p and pass HF_TOKEN if access is gated./results/train/checkpoints/iter_#########.pt file.Training writes full checkpoints under results/train/checkpoints/iter_#########.pt, updates results/train/checkpoints/latest_checkpoint.txt, and creates a cosmos_embed1_model_latest.pth symlink. For evaluate.checkpoint, inference.checkpoint, export.checkpoint, and train.resume_training_checkpoint_path, resolve and pass the exact iter_#########.pt file for the intended iteration. The action spec templates intentionally leave these checkpoint fields null so the model-skill runner or the user must provide the resolver-selected checkpoint. Use the latest symlink only when the user explicitly asks for latest.
For single-GPU resume/retrain from a consolidated checkpoint, set model.fsdp_shard_size: 1. The container default is 8, which sends resumed training through an FSDP apply path that Cosmos-Embed1 does not implement for this model class.
Variants:
| Variant | Resolution | Frames | Embedding dim |
|---|---|---|---|
Cosmos-Embed1-224p | 224 x 224 | 8 | 256 |
Cosmos-Embed1-336p | 336 x 336 | 8 | 768 |
Cosmos-Embed1-448p | 448 x 448 | 8 | 768 |
Keep model.network.embed_dim, model.input_hw, and model.network.spatial_resolution aligned with the selected variant.
| Parameter | Notes |
|---|---|
train.num_gpus | 1 for single GPU, >1 auto-launches torchrun, -1 auto-detects visible GPUs. |
train.max_iter | Main training length. Use 1 only for smoke testing. |
train.optim.optim | fused_adamw is faster when available; adamw is safer for smoke and portability. |
model.lora.enabled | Enables LoRA. Set model.network.visual_encoder.transformer_engine=false when LoRA is on. |
model.lora.lora_rank | LoRA rank. Start with 8; try 4, 8, or 16 for manual or AutoML-style sweeps. |
model.lora.lora_alpha | LoRA scaling factor. Start with 16; keep near 2 * lora_rank unless experiments show otherwise. |
model.lora.lora_dropout | LoRA dropout. Start with 0.1; sweep 0.0, 0.05, and 0.1 for small datasets. |
model.lora.bias | Bias policy: none, all, or lora_only. Keep none unless intentionally training biases. |
model.lora.use_rslora / use_dora | Optional LoRA variants. Enable one at a time and record the setting with the checkpoint. |
model.lora.target_modules | Optional module-name patterns for LoRA injection. Leave empty for the default ViT + Q-Former attention/MLP targets. |
model.lora.modules_to_save | Optional modules to keep fully trainable alongside LoRA. Leave empty unless preserving a task-specific head. |
evaluate.load_dataset_pkl / save_dataset_pkl | Cache evaluation embeddings. |
inference.load_dataset_pkl / save_dataset_pkl | Cache the search database for repeated retrieval. |
export.mode | video, text, combined, or huggingface. |
export.on_cpu | Recommended for export to avoid device mismatch issues. |
For parameter-efficient fine-tuning, set model.lora.enabled=true and keep
model.network.visual_encoder.transformer_engine=false; TAO Core's
Cosmos-Embed1 config notes that PEFT cannot inject adapters into Transformer
Engine layers. Treat the LoRA fields above as the first candidate parameters
for manual tuning or AutoML-style search before unfreezing larger model blocks.
Avoid changing target_modules or modules_to_save unless the user explicitly
needs custom adapter placement.
The Cosmos-Embed1 CLI consumes local paths and Python globs, not raw s3://.../*.mp4 URIs. For S3-backed runs, first stage a subset or full dataset to the execution host/container filesystem, then use local paths such as /data/video/*.mp4 in the spec.
Recommended S3 layout for staged MSR-VTT data:
s3://bucket/path/cosmos-embed/msrvtt-subset/
├── msrvtt_test_1k.json
└── video/
├── video7020.mp4
└── ...After downloading/syncing that prefix into the mounted data/ directory, use the same Docker commands above.
results/
├── train/
│ ├── cosmos_embed1_model_latest.pth
│ ├── cosmos_embed1_model_<iter>.pth
│ └── experiment.yaml
├── evaluate/
│ ├── metrics.json
│ └── experiment.yaml
├── inference/
│ ├── results.json
│ └── experiment.yaml
├── export/
│ ├── cosmos_embed1_combined.onnx
│ └── export_config.yaml
└── export_hf/
└── cosmos_embed1_hf/| Symptom | Cause | Fix |
|---|---|---|
MSRVTTDataset: 0 videos found | mp4_urls is not a local glob or metadata filenames do not match videos. | Mount data into the container and set mp4_urls=/data/video/*.mp4. |
| HF download/auth failure | Missing or invalid HF_TOKEN, or model agreement not accepted. | Accept the model terms and pass -e HF_TOKEN. |
cannot import name 'Imports' from 'wandb.proto.wandb_telemetry_pb2' | wandb==0.21.0 in the container is incompatible with protobuf==7.x. | Run python -m pip install "protobuf<7" in the container before invoking cosmos-embed1. |
Resume fails with Model does not implement 'apply_fsdp' | Single-GPU resume loaded a consolidated checkpoint while model.fsdp_shard_size stayed at the default 8. | Set model.fsdp_shard_size=1 for local single-GPU resume/retrain. |
| LoRA injection failure | Transformer Engine visual encoder is enabled. | Set model.network.visual_encoder.transformer_engine=false. |
| ONNX/HF export complains about missing components | Export checkpoint is partial or adapter-only. | Use a full checkpoint or configure pretrained visual/text sources before export. |
| CUDA OOM | Batch/resolution too high for the GPU. | Reduce batch size, use 224p, enable LoRA, or use more GPUs. |
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (references) in skills/tao-finetune-cosmos-embed of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Tao Finetune Cosmos Embed next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Finetune Cosmos Embed this skillNVIDIA/skills | 3.6k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | |
| Gemini Live APIgoogle/skills | 21k | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | |
| Swift Mlx Lmkellyvv/PhoneClaw | 1.3k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Civitaiartokun/comfyui-mcp | 803 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Wan T2v Videoartokun/comfyui-mcp | 803 | — | ~3.3k | Automated safety check: Pass | MIT | |
| Sft Launchopen-thoughts/OpenThoughts-Agent | 301 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 |
google/skills
Generates a Gemini LiveAPI client service class in the user's chosen programming language.
kellyvv/PhoneClaw
MLX Swift LM - Run LLMs and VLMs on Apple Silicon using MLX.
artokun/comfyui-mcp
Discover Civitai models with the BUILT-IN downloadmodel action:"searchcivitai" and install/generate them locally.
artokun/comfyui-mcp
Build WAN 2.2 Text-to-Video workflows. An agent skill from artokun/comfyui-mcp.
open-thoughts/OpenThoughts-Agent
Launch SFT via python -m hpc.launch --jobtype sft on any cluster (JSC Jupiter GH200, CINECA Leonardo A100, TACC Vista GH200), with EITHER backend — LLaMA-Factory (default) or axolotl (--sftbackend…
rand/cc-polymath
Automatically discover machine learning and AI skills when working with machine learning, PyTorch, training, inference, RAG, embeddings, fine-tuning, LLM, DSPy, HuggingFace, or diffusion models.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning. Tao Finetune Cosmos Embed is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Cosmos-Embed1 video-text embedding for text-to-video retrieval, video-to-video search, semantic deduplication, and fine-tuning.
Tao Finetune Cosmos Embed fits situations like: the user asks to fine-tune Cosmos-Embed1; run cosmos-embed inference; export Cosmos-Embed1; search videos with text.
Run `npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a claude-code`. Or copy the skill folder (skills/tao-finetune-cosmos-embed in NVIDIA/skills) into .claude/skills/tao-finetune-cosmos-embed in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a codex`. Or copy the skill folder (skills/tao-finetune-cosmos-embed in NVIDIA/skills) into .agents/skills/tao-finetune-cosmos-embed in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-finetune-cosmos-embed -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-finetune-cosmos-embed, .gemini/skills/tao-finetune-cosmos-embed, .github/skills/tao-finetune-cosmos-embed and .opencode/skills/tao-finetune-cosmos-embed in your project.
Going by SKILL.md and its folder, Tao Finetune Cosmos Embed needs the command-line tools its instructions call (docker, bash and python) and credentials named HF_TOKEN. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit, the published Cosmos-Embed TAO container (pinned in this skill), and a HuggingFace token when downloading pretrained `nvidia/Cosmos-Embed1-*` weights..
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Finetune Cosmos Embed is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Finetune Cosmos Embed: Gemini Live API (google/skills, 21k stars), Swift Mlx Lm (kellyvv/PhoneClaw, 1.3k stars), Civitai (artokun/comfyui-mcp, 803 stars) and Wan T2v Video (artokun/comfyui-mcp, 803 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.