Official agent skill

Tao Generate Video Reasoning Annotations

by NVIDIA in NVIDIA/skills

Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning…

OfficialApache-2.0Auto-check: notesMedia & Creative

Install Tao Generate Video Reasoning Annotations

skills CLI
$ npx skills add NVIDIA/skills --skill tao-generate-video-reasoning-annotations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-generate-video-reasoning-annotations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-generate-video-reasoning-annotations .claude/skills/tao-generate-video-reasoning-annotations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-generate-video-reasoning-annotations
GitHub stars
3.6k
Token cost
~2.8k tokens
SKILL.md length
1,003 words
Files
11 (incl. references)
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning…

  • Works in 5 steps: Videos → Domain — drives prompt selection → Anomaly / normal / mixed → …
  • The user wants to create video training data
  • SKILL.md covers Purpose, Pipeline architecture, Initial consultation and Quick start, plus 6 more sections
  • Runs Python scripts from its folder; needs GOOGLE_API_KEY

What it does

Tao Generate Video Reasoning Annotations is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning traces, via VLM/LLM distillation. Use when the user wants to "create video training data", "generate video QA datasets", "build CoT reasoning traces from videos", "auto-label videos", or run the videoreasoningannotation pipeline. Triggers include "video annotation", "video CoT", "video QA", "chain-of-thought", "video…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).

It sits in Media & Creative, covering AI video generation. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user wants to create video training data
  • Generate video QA datasets
  • Build CoT reasoning traces from videos
  • Auto-label videos

Example prompts

  • “create video training data”
  • “generate video QA datasets”
  • “build CoT reasoning traces from videos”
  • “/tao-generate-video-reasoning-annotations”

Requirements

  • Python 3
  • Docker
  • A credential in GOOGLE_API_KEY
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).
  • Pre-approved tools (allowed-tools): Read, Bash, Write

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Videos
  2. Domain — drives prompt selection
  3. Anomaly / normal / mixed
  4. VLM / LLM endpoint — confirm access before running
  5. Pilot vs full run

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Generate Video Reasoning Annotations loads about 2.8k tokens when it runs, and up to ~49k if it reads all its reference files. Until then it costs about 151 tokens; SKILL.md has 1,003 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~151
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~49k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 1,003 words, ~2,792 tokens.

Download SKILL.mdSave it as .claude/skills/tao-generate-video-reasoning-annotations/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
tao-generate-video-reasoning-annotations
description
Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning traces, via VLM/LLM distillation. Use when the user wants to "create video training data", "generate video QA datasets", "build CoT reasoning traces from videos", "auto-label videos", or run the video_reasoning_annotation pipeline. Triggers include "video annotation", "video CoT", "video QA", "chain-of-thought", "video captioning pipeline", "video distillation".
allowed-tools
Read, Bash, Write
compatibility
Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible).
license
Apache-2.0
metadata.author
NVIDIA Corporation
metadata.version
0.1.0
tags
video, annotation, chain-of-thought, captioning, qa-generation, vlm, llm, auto-label

Video Reasoning Annotation Pipeline

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Generate Chain-of-Thought training datasets from videos by producing multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with step-by-step reasoning traces. Domain-agnostic by default — customize prompts for any video domain.

Purpose

Transform raw videos into CoT Q&A training data for video understanding models. VLMs (e.g., Gemini, Qwen) act as "teacher" annotators: Steps 0–1 require the model to see the video (VLM calls); Steps 2–3 are text-to-text (cheaper LLM calls).

Pipeline architecture

Step 0:  [Optional] Filter & classify videos  → Keep domain-relevant, classify anomaly vs normal
Step 1a: Global + dense captions               → VLM: narrative summary + timestamped events
Step 1b: Chunk captions                         → VLM: fixed-duration segment micro-captions
Step 1c: [Optional, anomaly only] Highlight     → LLM extracts anomaly timestamp, VLM captions clip
Step 2:  Description synthesis                  → LLM: synthesize captions into structured narrative
Step 3:  QA generation                          → LLM: MCQ, binary, open-ended with reasoning
Step 4:  Parse outputs                          → Per-task `tao-vl-reason-v1.0` JSON files

Steps are individually selectable via workflow.steps. The pipeline has built-in resume — each step skips already-processed videos, so re-running after a prompt tweak is safe.

Initial consultation

When the user invokes this skill, walk through these questions in order. Don't skip — getting domain and VLM access right up front prevents wasted runs.

1. Videos
  • Path to the video directory and/or a JSONL with {"video_path": "..."} per line.
  • Confirm format (.mp4 preferred; .avi, .mov, .mkv also walked).
2. Domain — drives prompt selection

Ask the user: "What domain are these videos from?" Choose one of the following branches:

DomainWhat to do
generalUse the default prompts. Set prompts_module: "" (or omit). The built-in nvidia_tao_ds.auto_label.video_reasoning_annotation.prompts covers domain-agnostic content.
traffic (CCTV intersections, highways; dashcam excluded)Use the reference module. Set prompts_module: "nvidia_tao_ds.auto_label.video_reasoning_annotation.prompts_traffic", or copy references/prompts_traffic.py into the user's project and tune for their specific camera angles, then point prompts_module at the copy.
warehouse (industrial site CCTV — safety, operations, security)Same pattern. Set prompts_module: "nvidia_tao_ds.auto_label.video_reasoning_annotation.prompts_warehouse", or copy references/prompts_warehouse.py and tune.
custom (any other domain)Run the workshop in references/domain_adaptation.md. It walks through: Phase 1 — question types the user wants the model to answer; Phase 2 — caption-requirements checklist; Phase 3 — fill the [PLACEHOLDER] markers in nvidia_tao_ds.auto_label.video_reasoning_annotation.prompt_template. The two reference modules above are working examples to model after. Do this before any pipeline runs.
3. Anomaly / normal / mixed
  • Mixed dataset → workflow.mode: "auto" (Step 0 classifies each video).
  • Pre-split anomaly only → workflow.mode: "anomaly", drop Step 0.
  • Pre-split normal only → workflow.mode: "normal", drop Steps 0 and 1c.
4. VLM / LLM endpoint — confirm access before running
  • Gemini (default for both vlm.backend and llm.backend): user needs GOOGLE_API_KEY set, or to put the key in the YAML.
  • OpenAI-compatible (Qwen via vLLM, NIM endpoint, etc.): user provides base_url, model_name, and api_key.
  • Steps 2–3 are text-only — a smaller/cheaper LLM is fine for llm.backend even when vlm.backend is a frontier video model.

If the user has no endpoint at all and wants to self-host, point them at the skills/applications/tao-run-inference-service skill — a workflow that stands up a network-specific TAO inference microservice locally and exposes an OpenAI-compatible endpoint. Should support Cosmos, Qwen, and Gemma. Check skills/applications/tao-run-inference-service/references/service.yaml for the current valid_network_arch_config_basenames list before relying on a specific model.

If the user doesn't have endpoint access ready and isn't ready to set one up, stop here and help them figure it out first.

5. Pilot vs full run
  • Recommend a 5–10 video pilot when domain is custom, when any prompt was edited, or when this is the user's first run.
  • Full-run is fine for general / traffic / warehouse once the user has previously verified output quality on the same data type.
  • The pipeline has built-in resume, so a pilot followed by a full run does not re-process the pilot videos.

Quick start

The pipeline runs inside the TAO Toolkit container via the auto_label CLI:

bash
auto_label generate -e /path/to/spec.yaml \
    results_dir=/results \
    video_reasoning_annotation.data.video_root=/videos \
    video_reasoning_annotation.vlm.gemini.api_key=$GOOGLE_API_KEY \
    video_reasoning_annotation.workflow.mode=auto

Generate a default spec to start from:

bash
auto_label default_specs results_dir=/results module_name=auto_label
# then set:  autolabel_type: "video_reasoning_annotation"

All fields support Hydra dot-notation overrides on the command line. For the full YAML reference (every field, model/endpoint setup, error patterns), see references/configuration.md.

Show full SKILL.md (413 more words)Show less

Pilot workflow

Use this when running a 5–10 video pilot:

  1. Run the pipeline on the pilot subset with the chosen prompts_module and workflow.mode.
  2. Inspect results_dir/step_1a_caption/captions.jsonl — captions accurate, capturing the right level of detail?
  3. Inspect results_dir/step_3_qa/qa_output.jsonl — questions meaningful, answers correct, reasoning logical?
  4. If quality is insufficient: adjust the prompts (in prompts_module if domain-customized, or fall back to general if a domain module is over-tuned), and re-run. The pipeline auto-skips already-processed videos.
  5. Once satisfied, scale to the full dataset by pointing data.video_root (or data.input_jsonl_files) at the full set and re-running with the same results_dir (resume) or a fresh one (full re-run).

Quality compounds downstream — bad captions produce bad descriptions which produce bad QA. Focus iteration on Step 1a/1b output first; descriptions and QA usually improve once captions are right.

Configuration summary

Key fields (full reference in references/configuration.md):

FieldDefaultDescription
workflow.steps["0","1a","1b","1c","2","3","4"]Which pipeline steps to execute
workflow.mode"auto""auto", "anomaly", or "normal"
vlm.backend"gemini""gemini" or "openai" (OpenAI-compatible)
llm.backend"gemini"Same options; text-only, cheaper model works
workflow.max_workers4Parallel threads per step (watch API rate limits)
license""Optional: written to metadata.license in step 4 outputs (e.g. "CC-BY-4.0")
description_extra""Optional: extra text appended to per-task descriptions in step 4 metadata
prompts_module""Dotted import path to custom prompts module

Prompts

  • Built-in (general): nvidia_tao_ds.auto_label.video_reasoning_annotation.prompts — domain-agnostic, used by default.
  • Template: nvidia_tao_ds.auto_label.video_reasoning_annotation.prompt_template — same 26 keys with [PLACEHOLDER] markers for domain customization.
  • Reference modules (working examples for the consultation's traffic / warehouse branches): references/prompts_traffic.py, references/prompts_warehouse.py.
  • Custom domains: see references/domain_adaptation.md for the full workshop and placeholder reference.

Inputs

  • video_root: Directory of videos (walked recursively for .mp4, .avi, .mov, .mkv).
  • input_jsonl_files: List of JSONL files with {"video_path": "..."} per line. The video key is also accepted; extra fields are allowed.
  • filter_field: Optional boolean field to filter JSONL entries.

Provide video_root, input_jsonl_files, or both (lists merge).

Outputs

All outputs go to results_dir/ with per-step subdirectories (step_0_filter/, step_1a_caption/, …, step_4_output/):

  • Steps 0–3: JSONL — one JSON object per video per line.
  • Step 4: One <task>.json per non-empty task type, in the tao-vl-reason-v1.0 envelope. Up to 10 files: mcq.json, mcq_openended.json, bcq.json, bcq_openended.json, open_qa.json, causal_linkage.json, temporal_localization.json, temporal_description.json, scene_description.json, video_summarization.json.

Each step 4 file looks like:

json
{
  "format": "tao-vl-reason-v1.0",
  "metadata": {"type": "annotation", "task": "<task>", "date": "YYYY-MM-DD",
               "description": "<per-task + description_extra>", "license": "<from config>"},
  "media_root": "<data.video_root>" | null,
  "items": [{"video_id": "...", "question": "...", "answer": "...", "reasoning": "..."}, ...]
}

media_root mirrors data.video_root (or null when unset); each item's video_id is the entry's video path with the video_root prefix stripped. Set license and description_extra in the spec to populate the metadata.

Prerequisites

  • Container: nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt. <!-- versions-key: images.tao_toolkit.pyt -->
  • ffmpeg / ffprobe: required for chunk captioning (Step 1b) and highlight extraction (Step 1c).
  • VLM endpoint: at least one — Gemini API key or OpenAI-compatible endpoint.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/tao-generate-video-reasoning-annotations of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/configuration.md
  • references/domain_adaptation.md
  • references/prompts_traffic.py
  • references/prompts_warehouse.py
  • references/skill_info.yaml
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Tao Generate Video Reasoning Annotations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Generate Video Reasoning Annotations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Generate Video Reasoning Annotations this skillNVIDIA/skills3.6k—~2.8kAutomated safety check: NotesApache-2.0
Video Generationbytedance/deer-flow84k3 repos~1.4kAutomated safety check: PassMIT
Video Cover Imageitwanger/toBeBetterJavaer18k—~3.3kAutomated safety check: PassNone
Seedancesongguoxs/seedance-prompt-skill2.9k1 repos~2.5kAutomated safety check: PassNone
HyperFrames Video Entry Pointheygen-com/hyperframes60k3 repos~5.2kAutomated safety check: PassApache-2.0
Lanshu Create AI Presenter Videocclank/lanshu-create-ai-presenter-video2.6k—~3.6kAutomated safety check: PassMIT

Similar skills

  • Video Generation

    bytedance/deer-flow

    Generates short videos from a structured JSON prompt, optionally guided by a reference image used as the first or last frame.

    84k GitHub starsUsed in 3 repos~1.4k tokens
    Media & CreativeAuto-check passed
  • Video Cover Image

    itwanger/toBeBetterJavaer

    Generate matched 3:4, 16:9, and 4:3 short-video cover images from toBeBetterJavaer video scripts or AI/Java technical topics.

    18k GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Seedance

    songguoxs/seedance-prompt-skill

    This skill should be used when the user asks to "generate video prompts", "create Seedance prompts", "write video descriptions", mentions "Seedance", "seedance", "即梦", "即梦平台", "视频提示词", "视频生成"…

    2.9k GitHub starsUsed in 1 repo~2.5k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Video Entry Point

    heygen-com/hyperframes

    Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.

    60k GitHub starsUsed in 3 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Lanshu Create AI Presenter Video

    cclank/lanshu-create-ai-presenter-video

    Turn a topic or finished script into a complete, publish-ready explainer video — led by an AI presenter from an authorized adult presenter image, or performed in one of nine visual explainer styles…

    2.6k GitHub stars~3.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Video Shots

    eternityspring/reelbench-skills

    拉片:把一条成片拆成逐镜头的分析表——每个镜头的时长、景别、类别、运镜、画面. An agent skill from eternityspring/reelbench-skills.

    878 GitHub starsUsed in 1 repo~1.8k tokens
    Media & CreativeAuto-check: notes

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Tao Generate Video Reasoning Annotations

What does Tao Generate Video Reasoning Annotations do?

Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning…. Tao Generate Video Reasoning Annotations is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Multi-step video annotation pipeline that turns raw videos into Chain-of-Thought training data — multi-level captions, structured descriptions, and QA pairs (MCQ, binary, open-ended) with reasoning traces, via VLM/LLM distillation.

When should I use Tao Generate Video Reasoning Annotations?

Tao Generate Video Reasoning Annotations fits situations like: the user wants to create video training data; generate video QA datasets; build CoT reasoning traces from videos; auto-label videos.

How do I install Tao Generate Video Reasoning Annotations in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-generate-video-reasoning-annotations -a claude-code`. Or copy the skill folder (skills/tao-generate-video-reasoning-annotations in NVIDIA/skills) into .claude/skills/tao-generate-video-reasoning-annotations in your project. Claude Code loads it when a task matches its description.

How do I install Tao Generate Video Reasoning Annotations in Codex?

Run `npx skills add NVIDIA/skills --skill tao-generate-video-reasoning-annotations -a codex`. Or copy the skill folder (skills/tao-generate-video-reasoning-annotations in NVIDIA/skills) into .agents/skills/tao-generate-video-reasoning-annotations in your project. Codex loads it when a task matches its description.

Can I use Tao Generate Video Reasoning Annotations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-generate-video-reasoning-annotations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-generate-video-reasoning-annotations, .gemini/skills/tao-generate-video-reasoning-annotations, .github/skills/tao-generate-video-reasoning-annotations and .opencode/skills/tao-generate-video-reasoning-annotations in your project.

What does Tao Generate Video Reasoning Annotations need to run?

Going by SKILL.md and its folder, Tao Generate Video Reasoning Annotations needs Python for the scripts in its folder and credentials named GOOGLE_API_KEY. Our summary lists: Python 3; Docker; A credential in GOOGLE_API_KEY. Its frontmatter pre-approves these tools: Read, Bash, Write. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit + at least one VLM endpoint (Gemini API key or OpenAI-compatible)..

Does Tao Generate Video Reasoning Annotations access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Tao Generate Video Reasoning Annotations safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Generate Video Reasoning Annotations use?

Tao Generate Video Reasoning Annotations is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Generate Video Reasoning Annotations use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 47k tokens, read only when the agent opens those files.

What are the alternatives to Tao Generate Video Reasoning Annotations?

Skills that share tags, products or a category with Tao Generate Video Reasoning Annotations: Video Generation (bytedance/deer-flow, 84k stars), Video Cover Image (itwanger/toBeBetterJavaer, 18k stars), Seedance (songguoxs/seedance-prompt-skill, 2.9k stars) and HyperFrames Video Entry Point (heygen-com/hyperframes, 60k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Generate Video Reasoning Annotations?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.