Official agent skill

Tao Train Action Recognition

by NVIDIA in NVIDIA/skills

Action recognition from video sequences. An agent skill from NVIDIA/skills.

OfficialApache-2.0Auto-check: notesMedia & Creative

Install Tao Train Action Recognition

skills CLI
$ npx skills add NVIDIA/skills --skill tao-train-action-recognition -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills tao-train-action-recognition --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-action-recognition .claude/skills/tao-train-action-recognition && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tao-train-action-recognition
GitHub stars
3.5k
Token cost
~3k tokens
SKILL.md length
1,104 words
Files
16 (incl. references)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Action recognition from video sequences. An agent skill from NVIDIA/skills.

  • Running inference on a TAO action-recognition model
  • SKILL.md covers Quick Start (docker run), Dataclass Schemas, Train Action Policy and Training Requirements, plus 6 more sections
  • Calls docker
  • Phrases include train action recognition

What it does

Tao Train Action Recognition is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Action recognition from video sequences. Supports RGB, optical flow, and joint (multi-stream) input types for classifying temporal actions in video clips. Use when training, evaluating, exporting, or running inference on a TAO action-recognition model. Trigger phrases include "train action recognition", "video action classification", "RGB + optical flow action model", "TAO ActionRecognition".

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.

It sits in Media & Creative, covering Video production. It works with Docker. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Running inference on a TAO action-recognition model
  • Phrases include train action recognition
  • Video action classification
  • RGB + optical flow action model

Example prompts

  • “train action recognition”
  • “video action classification”
  • “RGB + optical flow action model”
  • “/tao-train-action-recognition”

Requirements

  • Python 3
  • Docker
  • Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit.
  • Pre-approved tools (allowed-tools): Read, Bash

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Bash

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker + nvidia-container-toolkit.

    From compatibility in the SKILL.md frontmatter.

Context cost

Tao Train Action Recognition loads about 3k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 1,104 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Bash

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,104 words, ~2,966 tokens.

Download SKILL.mdSave it as .claude/skills/tao-train-action-recognition/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
tao-train-action-recognition
description
Action recognition from video sequences. Supports RGB, optical flow, and joint (multi-stream) input types for classifying temporal actions in video clips. Use when training, evaluating, exporting, or running inference on a TAO action-recognition model. Trigger phrases include "train action recognition", "video action classification", "RGB + optical flow action model", "TAO ActionRecognition".
allowed-tools
Read, Bash
compatibility
Requires docker + nvidia-container-toolkit.
license
Apache-2.0
metadata.version
0.1.0
metadata.author
NVIDIA Corporation
tags
action, recognition

Action Recognition

Standalone install? If this session was not initialized by the TAO skill bank plugin, run the tao-setup skill first (host preflight, credentials, cross-skill discovery).

Action recognition from video sequences. Supports RGB, optical flow, and joint (multi-stream) input types for classifying temporal actions in video clips.

Set model.pretrained_model_path for pretrained backbone weights.

Quick Start (docker run)

Docker-native launch — no TAO SDK and no Python on the host. Use the local Docker/platform skill instead when it gives a stricter environment-specific command (non-root UID mapping, cache redirects, remote daemons).

bash
TAO_PYT_IMAGE_DEFAULT=nvcr.io/nvidia/tao/tao-toolkit:7.2.0-pyt  # versions-key: images.tao_toolkit.pyt
TAO_PYT_IMAGE="${TAO_PYT_IMAGE:-$TAO_PYT_IMAGE_DEFAULT}"
RUN_ROOT="${RUN_ROOT:-$PWD}"
DOCKER_COMMON=(
  --rm --gpus all --shm-size=8g
  --shm-size=8g
  --ulimit memlock=-1
  --ulimit stack=67108864
  -v "$RUN_ROOT/data:/data:ro"
  -v "$RUN_ROOT/specs:/specs:ro"
  -v "$RUN_ROOT/results:/results"
)

Train:

bash
docker run "${DOCKER_COMMON[@]}" "$TAO_PYT_IMAGE" \
  action_recognition train -e /specs/train.yaml

Evaluate:

bash
docker run "${DOCKER_COMMON[@]}" "$TAO_PYT_IMAGE" \
  action_recognition evaluate -e /specs/evaluate.yaml

Inference:

bash
docker run "${DOCKER_COMMON[@]}" "$TAO_PYT_IMAGE" \
  action_recognition inference -e /specs/inference.yaml

Export:

bash
docker run "${DOCKER_COMMON[@]}" "$TAO_PYT_IMAGE" \
  action_recognition export -e /specs/export.yaml

Every action takes its spec with -e; results_dir is set in the spec or overridden on the command line. Mount any pretrained-weights directory the spec references, and keep every in-container path consistent across actions.

Dataclass Schemas

Generated TAO Core schemas are packaged in schemas/<action>.schema.json, with schemas/manifest.json listing available actions. Each generated schema also emits references/spec_template_<action>.yaml from the schema top-level default field. AutoML enablement is declared at the model layer in references/skill_info.yaml via automl_enabled. Runnable AutoML for an action requires schemas/<action>.schema.json and references/spec_template_<action>.yaml to exist and parse. Use the packaged selected-action schema for automl_default_parameters, automl_disabled_parameters, defaults, min/max bounds, enums, option weights, math conditions, dependencies, and popular parameters. Do not expect ~/tao-core at runtime; maintainers regenerate schemas/templates before packaging the skill bank.

Train Action Policy

This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.

Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.

Training Requirements

  • Dataset type: action_recognition
  • Formats: default
  • Training monitoring metrics: val_loss, val_acc
  • Evaluate task metrics: accuracy, m_accuracy
  • AutoML metric contract: for the required evaluation-backed baseline and final comparison, use accuracy with maximize direction and run evaluate through eval_fn for every recommendation. val_loss is suitable only for an explicitly accepted training-proxy run because evaluate does not emit it. Scratch runs with no starting checkpoint require a minimal default train job followed by evaluation of its exact epoch/step checkpoint for the baseline.
Per-Action Dataset Requirements
ActionSpec KeySourceFilesList?
evaluateevaluate.test_dataset_dirtrain_datasetstest/ extracted from test.tar.gzNo
inferenceinference.inference_dataset_dirtrain_datasetstest/smile/ extracted from test/smile.tar.gzNo
traindataset.train_dataset_dirtrain_datasetstrain/ extracted from train.tar.gzNo
traindataset.val_dataset_dirtrain_datasetstest/ extracted from test.tar.gzNo
Typical Spec Overrides

Data source overrides are mandatory for every action — the agent MUST construct data source paths from the Per-Action Dataset Requirements table above and include them in spec_overrides.

python
LOCAL_DATA = "/workspace/data/extracted"

If the source dataset is provided as the TAO sample archives train.tar.gz, test.tar.gz, or test/smile.tar.gz, download and extract them before launching the TAO container. The action-recognition entrypoints expect directory paths and fail with NotADirectoryError when these spec keys point at .tar.gz files.

train (mandatory data sources):

python
{
    "train.num_epochs": 30,
    "train.checkpoint_interval": 10,
    "train.validation_interval": 10,
    "train.num_gpus": 1,
    "dataset.label_map": {
        "catch": 0,
        "smile": 1
    },
    "dataset.batch_size": 2,
    "dataset.train_dataset_dir": f"{LOCAL_DATA}/train",
    "dataset.val_dataset_dir": f"{LOCAL_DATA}/test",
}

evaluate (mandatory data sources):

python
{
    "dataset.label_map": {
        "catch": 0,
        "smile": 1
    },
    "evaluate.test_dataset_dir": f"{LOCAL_DATA}/test",
}

inference (mandatory data sources):

python
{
    "dataset.label_map": {
        "catch": 0,
        "smile": 1
    },
    "inference.inference_dataset_dir": f"{LOCAL_DATA}/smile_infer/smile",
}

export (mandatory checkpoint + output path):

python
{
    "export.checkpoint": "<selected train checkpoint>",
    "export.onnx_file": "<results_dir>/action_recognition.onnx",
}

For direct local-docker chaining without the SDK resolver, select the concrete checkpoint produced by training, for example model_epoch_000_step_00005.pth, and pass that exact file to evaluate, inference, and export. Do not use the ar_model_latest.pth symlink unless the user explicitly requests latest-checkpoint behavior. For resume training, set train.resume_training_checkpoint_path to the exact epoch/step checkpoint being resumed.

Show full SKILL.md (465 more words)Show less

Eval Dataset

Optional. Test dataset may be distributed as test.tar.gz separate from training; extract it and point the spec to the extracted test/ directory. TAO training emits val_loss and val_acc for the packaged sample data, while the evaluate action emits accuracy and m_accuracy. Use accuracy with maximize direction for the normal evaluation-backed AutoML workflow. Use val_loss with minimize direction only when the user explicitly accepts a training-only proxy without the required impact baseline.

Important Parameters

  • model.model_type: Input type: rgb, of (optical flow), or joint (multi-stream).
  • model.backbone: Default resnet_18. Used as the spatial feature extractor.
  • dataset.label_map: Dictionary mapping class names to indices.
  • model.rgb_seq_length: Number of frames per clip for RGB input.
  • model.of_seq_length: Number of frames for optical flow input.
  • train.optim.lr: Learning rate. Default 5e-4.

Multi-GPU / Multi-Node

Launch method: Lightning-managed (single python process, Lightning spawns workers).

Spec KeyDescriptionDefault
train.num_gpusNumber of GPUs1
train.gpu_idsGPU device indices[0]
  • Strategy: auto (Lightning picks best strategy automatically)
  • No explicit num_nodes or distributed_strategy config — single-node oriented

Hardware

Minimum 1 GPU(s), recommended 2 GPU(s). 16GB+ VRAM per GPU. Memory depends on sequence length and input resolution. batch_size=2 is conservative for video data.

Error Patterns

Sequence length mismatch: Ensure video clips have enough frames for the configured rgb_seq_length or of_seq_length.

Evaluate/inference missing label map: Downstream actions rebuild the ActionRecognitionModel before loading the checkpoint, so they need the same dataset.label_map used during training. Include it with every evaluate or inference spec; otherwise model construction fails before the checkpoint can be validated.

Spec Param / Parent Model Inference

Model-specific inference mappings belong in this MD file, not in config.json. Generated runners should read this section and apply the mappings with SDK helpers before create_job(). This mirrors the old microservices infer_params.py flow.

Inference mappings from TAO Core action_recognition.config.json:

ActionSpec FieldInference FunctionMeaning
evaluateencryption_keykeyencryption key
evaluateevaluate.checkpointparent_modelmodel file inferred from the parent job results folder
evaluateresults_diroutput_dircurrent job results directory
exportencryption_keykeyencryption key
exportexport.checkpointparent_modelmodel file inferred from the parent job results folder
exportexport.onnx_filecreate_onnx_fileoutput ONNX path
exportresults_diroutput_dircurrent job results directory
inferenceencryption_keykeyencryption key
inferenceinference.checkpointparent_modelmodel file inferred from the parent job results folder
inferenceresults_diroutput_dircurrent job results directory
trainencryption_keykeyencryption key
trainmodel.of_pretrained_model_pathptm_if_no_resume_modelPTM when no resume checkpoint exists
trainmodel.rgb_pretrained_model_pathptm_if_no_resume_modelPTM when no resume checkpoint exists
trainresults_diroutput_dircurrent job results directory
traintrain.resume_training_checkpoint_pathresume_modelmodel file inferred from the current job results folder

For parent_model or parent_model_folder, pass the upstream train/export/AutoML child job id as parent_job_id. The SDK lists the parent result folder, filters checkpoint artifacts, and returns the selected model file or folder. Do not add these mappings back to config.json and do not patch generated runner scripts to guess checkpoint paths.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (references) in skills/tao-train-action-recognition of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/skillspector-baseline.yaml
  • evals/evals.json
  • references/skill_info.yaml
  • references/spec_template_evaluate.yaml
  • references/spec_template_export.yaml
  • references/spec_template_inference.yaml
  • references/spec_template_train.yaml
  • schemas/evaluate.schema.json
  • schemas/export.schema.json
  • schemas/inference.schema.json
  • schemas/manifest.json
  • schemas/train.schema.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Tao Train Action Recognition next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Tao Train Action Recognition compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Tao Train Action Recognition this skillNVIDIA/skills3.5k—~3kAutomated safety check: NotesApache-2.0
Vhs Demobabarot/gh-infra136—~1.2kAutomated safety check: PassMIT
Setup Mulmoclaudereceptron/mulmoclaude368—~1kAutomated safety check: NotesMIT
Floor Bot Analysisfossasia/eventyay-interpretation1.6k—~1kAutomated safety check: NotesApache-2.0
MediaGo Video Downloadermediago-dev/mediago9.3k—~1.1kAutomated safety check: PassMIT
Video Clip Repurposingpawbytes/skill-suites113—~1.2kAutomated safety check: PassMIT

Similar skills

  • Vhs Demo

    babarot/gh-infra

    A skill your agent uses when running demo recordings, diagnosing recording failures, or regenerating GIFs from existing MP4s.

    136 GitHub stars~1.2k tokensUpdated 10 days ago
    DevOps & CloudAuto-check passed
  • Setup Mulmoclaude

    receptron/mulmoclaude

    Interactively guide MulmoClaude setup following README instructions.

    368 GitHub stars~1k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Floor Bot Analysis

    fossasia/eventyay-interpretation

    Use this skill for floor bot analysis.

    1.6k GitHub stars~1k tokensUpdated 4 days ago
    Media & CreativeAuto-check: notes
  • MediaGo Video Downloader

    mediago-dev/mediago

    Downloads videos from m3u8/HLS streams, Bilibili and direct URLs by driving a running MediaGo instance's REST API, with bilingual first-time setup guidance.

    9.3k GitHub stars~1.1k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Video Clip Repurposing

    pawbytes/skill-suites

    Cuts long-form video into short platform-ready clips: finds strong moments, reframes for vertical, adds subtitles and brand overlays, and writes a clip manifest.

    113 GitHub stars~1.2k tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • Runpod

    digitalsamba/claude-code-video-toolkit

    Cloud GPU processing via RunPod serverless. An agent skill from digitalsamba/claude-code-video-toolkit.

    2.2k GitHub stars~2.1k tokensUpdated 4 days ago
    Backend & APIsAuto-check: notes

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Works with

Questions about Tao Train Action Recognition

What does Tao Train Action Recognition do?

Action recognition from video sequences. An agent skill from NVIDIA/skills. Tao Train Action Recognition is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Action recognition from video sequences.

When should I use Tao Train Action Recognition?

Tao Train Action Recognition fits situations like: running inference on a TAO action-recognition model; phrases include train action recognition; video action classification; RGB + optical flow action model.

How do I install Tao Train Action Recognition in Claude Code?

Run `npx skills add NVIDIA/skills --skill tao-train-action-recognition -a claude-code`. Or copy the skill folder (skills/tao-train-action-recognition in NVIDIA/skills) into .claude/skills/tao-train-action-recognition in your project. Claude Code loads it when a task matches its description.

How do I install Tao Train Action Recognition in Codex?

Run `npx skills add NVIDIA/skills --skill tao-train-action-recognition -a codex`. Or copy the skill folder (skills/tao-train-action-recognition in NVIDIA/skills) into .agents/skills/tao-train-action-recognition in your project. Codex loads it when a task matches its description.

Can I use Tao Train Action Recognition in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-action-recognition -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-action-recognition, .gemini/skills/tao-train-action-recognition, .github/skills/tao-train-action-recognition and .opencode/skills/tao-train-action-recognition in your project.

What does Tao Train Action Recognition need to run?

Going by SKILL.md and its folder, Tao Train Action Recognition needs the command-line tools its instructions call (docker). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..

Does Tao Train Action Recognition access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Tao Train Action Recognition safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Tao Train Action Recognition use?

Tao Train Action Recognition is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Tao Train Action Recognition use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Tao Train Action Recognition?

Skills that share tags, products or a category with Tao Train Action Recognition: Vhs Demo (babarot/gh-infra, 136 stars), Setup Mulmoclaude (receptron/mulmoclaude, 368 stars), Floor Bot Analysis (fossasia/eventyay-interpretation, 1.6k stars) and MediaGo Video Downloader (mediago-dev/mediago, 9.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Tao Train Action Recognition?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.