Official agent skill

Physical AI Video Augmentation on OSMO

by NVIDIA in NVIDIA/skills

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install Physical AI Video Augmentation on OSMO

skills CLI
$ npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills physical-ai-video-data-augmentation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/physical-ai-video-data-augmentation .claude/skills/physical-ai-video-data-augmentation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
physical-ai-video-data-augmentation
GitHub stars
3.5k
Token cost
~4.7k tokens
SKILL.md length
1,809 words
Files
103 (incl. scripts, references, assets)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

  • Works in 6 steps: Select the workflow (auto_labeling,… → Provide a tentative execution-time… → Run preflight and readiness checks… → …
  • Running a video data augmentation job on an OSMO cluster
  • SKILL.md covers Purpose, Prerequisites, Instructions and Available Scripts, plus 9 more sections
  • Calls bash, git and python3; needs NGC_API_KEY and NGC_CLI_API_KEY

What it does

The agent picks one of four workflows from your intent (`auto_labeling`, `augmentation_and_al`, `e2e` or `e2e_super_resolution`), gives a rough execution-time overview, and runs preflight and readiness checks before anything is submitted. Submit-time values come from the active dataset backend, and it must never guess `storage_url`. It then submits with explicit interpolation values, monitors to completion, retrieves outputs and provides side-by-side comparison evidence for augmented flows.

Prerequisites are checked first: an optional NGC API key, a Hugging Face token for gated Cosmos and SeedVR weights, the `osmo` CLI logged in with a default profile and a matching data credential profile, and at least one online GPU pool. Missing secrets surface as `USER_INPUT_REQUIRED` from `scripts/preflight_credentials.sh`. The folder ships 109 files, including OSMO configs and cookbooks. Questions about tuning inside the container are out of scope.

When your agent uses it

  • Running a video data augmentation job on an OSMO cluster
  • Auto-labeling a set of videos with pseudo labels
  • Checking credentials and GPU pool readiness before submitting a workflow
  • Downloading outputs and comparing augmented clips with the originals

Example prompts

  • “Run the end-to-end video augmentation workflow on my city traffic clips in OSMO.”
  • “Do the preflight checks for an auto labeling run and tell me what credentials are missing.”
  • “Monitor the submitted workflow and download the outputs when it finishes.”

Requirements

  • The `osmo` CLI, logged in with a default profile and data credential profile
  • A Hugging Face token for gated model weights
  • At least one online GPU pool on OSMO

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Select the workflow (auto_labeling, augmentation_and_al, e2e,
  2. Provide a tentative execution-time overview before starting run actions.
  3. Run preflight and readiness checks before submit.
  4. Derive submit-time values from the active dataset backend (never guess
  5. Submit the workflow with explicit interpolation values and monitor to completion.
  6. Retrieve outputs, provide side-by-side comparison evidence for augmented

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • git
    • python3
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NGC_API_KEY
    • NGC_CLI_API_KEY
    • NVIDIA_API_KEY
    • OPENAI_API_KEY
    • VLM_API_KEY
    • LLM_API_KEY
    • HF_TOKEN
    • HUGGING_FACE_HUB_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Physical AI Video Augmentation on OSMO loads about 4.7k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 1,809 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:234
    light does not require a workload-local `.env`. Runtime interpolation is

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,809 words, ~4,747 tokens.

Download SKILL.mdSave it as .claude/skills/physical-ai-video-data-augmentation/SKILL.md (or your agent's skills folder). This skill also uses 102 other files; get the full folder from GitHub.
name
physical-ai-video-data-augmentation
description
Use when running video data augmentation and auto-labeling workflows on OSMO: flow selection, preflight, submit-time interpolation, monitoring, and output retrieval. Trigger keywords: video data augmentation, data enrichment, auto labeling, VDA demo, OSMO workflow, pseudo labeling.
license
CC-BY-4.0 AND Apache-2.0
metadata.owner
NVIDIA
metadata.service
data
metadata.version
1.0.0
metadata.reviewed
2026-05-26
metadata.author
NVIDIA
metadata.tags
physical-ai, video-data-augmentation, auto-labeling, cosmos

Physical AI Video Data Augmentation Workflow Orchestrator

Default workflow skill for VDA execution on OSMO. It owns flow selection, preflight, cache readiness, inference-path decisions, submit-time interpolation, monitoring, and output retrieval. Component skills are consult-only.

Purpose

Run the end-to-end VDA workflow safely and reproducibly from preflight to output download.

Do NOT use this skill for container-internal tuning-only questions.

Prerequisites

Confirm these before running preflight or any submit. Missing required secrets surface as USER_INPUT_REQUIRED: from scripts/preflight_credentials.sh.

RequirementHow it is satisfiedUsed for
NGC API key (optional)NGC_API_KEY, NGC_CLI_API_KEY, or compatible nvapi-* token in NVIDIA_API_KEY/OPENAI_API_KEY/VLM_API_KEY/LLM_API_KEYOptional for nvcr_io credential refresh and NGC REST scope probe; default VDA image refs are validated via workflow registry probes
Hugging Face tokenHF_TOKEN (or HUGGING_FACE_HUB_TOKEN), or a cached token at ~/.cache/huggingface/tokenCreates the OSMO hf_token credential; pulls gated Cosmos/SeedVR weights
OSMO CLI accessosmo on PATH, logged in, with a default profile and a registered DATA credential profile matching storage_urlSubmitting/monitoring workflows and listing/downloading objects
GPU poolAt least one ONLINE pool in osmo pool list --mode free; POD_TEMPLATE carries GPU toleration/selectorsScheduling setup + worker tasks

Optional (only for the strict NGC org/team probe): NGC_ORG + NGC_TEAM (or NGC_CLI_ORG / NGC_CLI_TEAM). External VLM/LLM endpoint keys are validated separately, not by preflight.

Key handling rule: nvapi-* tokens are first-class inputs for nvcr_io. Never reject by token prefix alone; use workflow registry probe results as source of truth.

Instructions

  1. Select the workflow (auto_labeling, augmentation_and_al, e2e, e2e_super_resolution) from user intent.
  2. Provide a tentative execution-time overview before starting run actions.
  3. Run preflight and readiness checks before submit.
  4. Derive submit-time values from the active dataset backend (never guess storage_url).
  5. Submit the workflow with explicit interpolation values and monitor to completion.
  6. Retrieve outputs, provide side-by-side comparison evidence for augmented flows, and summarize task outcomes.

Use run_script(...) for script execution. Canonical examples:

python
run_script("bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/augmentation_and_al.yaml")
run_script("python3 scripts/pre_submit_guard.py --workflow assets/configs/osmo/auto_labeling.yaml")
run_script("bash scripts/prepare_demo_assets.sh /srv/sdg/data/vda_inputs")

Available Scripts

Use script-level --help for exact arguments.

ScriptRole
scripts/preflight_credentials.shSecrets/control-plane preflight and workflow image access checks
scripts/pre_submit_guard.pySubmit-time interpolation, cache, and dataset safety checks
scripts/prepare_demo_assets.shDemo video pull + flatten for default demo path
scripts/generate_configs.pySetup-time config and cookbook projection generation
scripts/cosmos_worker.shAugmentation worker execution
scripts/pl_original_worker.shOriginal-video auto-labeling worker execution
scripts/pl_augmented_worker.shAugmented-video auto-labeling worker execution
scripts/osmo_barrier.pyMulti-node barrier synchronization
scripts/stage_run_artifacts.shLocal mirror of full run output + input video
scripts/render_side_by_side.shSide-by-side comparison render from local artifacts

Supported Flows

FlowOSMO YAMLGroup sequenceTypical use
augmentation_and_alassets/configs/osmo/augmentation_and_al.yamlsetup -> augmentation -> auto_labeling_augmentedAugment one or more videos, then auto-label augmented outputs
auto_labelingassets/configs/osmo/auto_labeling.yamlsetup -> auto_labelingLabel original videos only
e2eassets/configs/osmo/e2e.yamlsetup -> (auto_labeling_original + augmentation) -> auto_labeling_augmentedThroughput-first path
e2e_super_resolutionassets/configs/osmo/e2e_super_resolution.yamlsetup -> auto_labeling_original -> augmentation -> auto_labeling_augmentedSequential path with SR gate before augmentation

Legacy alias assets/configs/osmo/augmentation_and_pl.yaml remains for backwards compatibility.

Pick the right workflow for the user's request
User intentWorkflow
"Label my source videos" / "PL-only" / "no augmentation"auto_labeling
"Create augmented videos and label them"augmentation_and_al
"Run the full pipeline quickly"e2e
"Run full pipeline, but gate on SR-enhanced originals first"e2e_super_resolution

Disambiguation: handle vague requests before committing

Default to autonomy: ask only when missing information blocks execution.

Autonomous defaults (do NOT ask)
  • If dataset source is absent, run VDA demo path (scripts/prepare_demo_assets.sh) and continue with dataset=vda-demo.
  • If flow is not explicitly requested, default to augmentation_and_al.
  • If endpoint mode is unspecified, default to in-cluster persistent NIM reuse and automatic NIM deploy/repair when unhealthy.
  • If cache is missing, run setup_model_cache.yaml, rerun pre-submit guard, and continue automatically on success.
  • After any stage completes successfully, continue to the next stage immediately. Do not pause with "Ready when you are" or equivalent approval prompts.
Triggers that should pause for disambiguation
Missing inputWhy it mattersAsk
USER_INPUT_REQUIRED from preflightRequired secret is missingAsk one concise unblock question for exactly the missing value(s)
Storage backend prefix cannot be derived from the active dataset/upload rootWrong scheme causes runtime storage auth mismatch"What is the backend-native root prefix for this run?"
No ONLINE GPU pool/platform can be selectedWorkflow cannot schedule setup/workers"Which GPU pool/platform should this run target?"
When NOT to disambiguate
  • Do not ask for cookbook unless user explicitly asks to change scene profile.
  • Do not offer external endpoints by default.
  • Do not ask A/B cache strategy questions; default is automatic cache setup.
  • Do not ask to scale down existing NIMs; this is forbidden.
  • Do not invent, scrape, or generate random videos when input is missing.
  • Do not use non-VDA demo sources (for example Carline adaptation assets) unless the user explicitly requests a different dataset.

Step 0: Select Flow and Gather Inputs

Input video policy (non-negotiable)
  • Always preserve user-provided video inputs (dataset URL, local path, or upload folder) as first-class and preferred.
  • Never replace an explicit user video with demo assets or any other source.
  • If no video input is provided, default to VDA demo assets via scripts/prepare_demo_assets.sh (HF dataset flow) without asking extra source-selection questions.
  • If the user explicitly mentions an input video or dataset, prefer and use that input instead of demo assets.
  • Use only VDA demo assets (nvidia/video-data-augmentation-demo) for the default demo path.
  • Never propose arbitrary web clip downloads or placeholder videos unless the user explicitly requests that behavior.

Collect only missing values:

  1. Dataset source (prefer explicit user-provided dataset_url or local upload folder; otherwise default to VDA demo assets and proceed).
  2. Flow (auto_labeling, augmentation_and_al, e2e, e2e_super_resolution); default to augmentation_and_al when unspecified.
  3. OSMO gpu_platform for all VDA resources (auto-select an ONLINE platform when unambiguous; ask only when no valid option exists).
  4. Endpoint mode (default in-cluster NIM reuse/deploy unless explicitly overridden).

Do not guess gpu_platform (for example microk8s). Use the exact current platform label shown by osmo pool list --mode free (for example gpu).

Generate run stamp before each submit:

bash
STAMP=$(cat /proc/sys/kernel/random/uuid | cut -c1-8)
RUN_ID="run-$STAMP"

Execution Time Overview (required before run)

Before running any mutating command (osmo credential set, NIM install/repair, cache workflow submit, or target VDA workflow submit), provide a short ETA overview to the user.

Keep it concise (one short paragraph or 4-6 bullets) and include:

  • whether this looks like a cold start (NIM/cache missing) or warm start (NIM/cache already healthy),
  • major phases with approximate durations,
  • a total expected range for the selected workflow.

Baseline ranges (from observed MicroK8s + OSMO runs):

PhaseTypical duration
Credentials + preflight~1-2 min
NIM deploy/download/warmup (if needed)~10-15 min
Demo assets download/upload (if demo path)~1-3 min
Model cache population (if needed)~15-25 min
Workflow submit + queue/start~1-3 min

Workflow runtime ranges after submit:

FlowTypical runtime
auto_labeling~6-15 min
augmentation_and_al~20-35 min
e2e~22-40 min
e2e_super_resolution~25-45 min

Cold-start end-to-end runs are commonly ~45-80 min; warm-start runs are usually ~20-45 min depending on flow and video length.

Show full SKILL.md (797 more words)Show less

Common Preconditions (all flows)

  1. Credential and control-plane preflight

    bash
    bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/<mode>.yaml

    Restricted egress:

    bash
    bash scripts/preflight_credentials.sh --no-probe --workflow assets/configs/osmo/<mode>.yaml

    Preflight does not require a workload-local .env. Runtime interpolation is driven by submit-time values (dataset, run_id, gpu_platform, video, storage_url, skills_dir) supplied in one --set-string list.

    Passing --workflow validates pull access for the active workflow image refs (workflow.groups[].tasks[].image) using anonymous bearer access with credential fallback when provided. If replacement NGC/HF secrets are provided in env, preflight refreshes existing nvcr_io / hf_token automatically when present. Use --refresh to force overwrite even when no new env secrets were supplied:

    bash
    bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/<mode>.yaml --refresh

    If output contains USER_INPUT_REQUIRED:, ask one concise unblock question and stop.

    On workflow image 401/403, report registry access failure after probe checks on the listed image refs; do not claim a key family (for example nvapi-*) is categorically unsupported.

  2. Storage interpolation policy

    storage_url must be derived from the actual dataset/upload backend for the current run.

    text
    dataset_url=azure://storiondevxah69/osmo-workflows/datasets/vda-demo
    storage_url=azure://storiondevxah69/osmo-workflows
    dataset=vda-demo

    Never silently default to stale s3:// values on non-S3 backends.

  3. Inference policy (non-negotiable)

    • Reuse healthy in-cluster persistent NIM endpoints by default.
    • If missing/unhealthy, deploy automatically — this is a prerequisite, not a user decision. Do NOT pause to ask; run the install with the VDA allow-list:
    bash
    export NIM_SERVICES="qwen3-vl qwen25-14b"
    skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/inference-nim-operator/scripts/install.sh
    • See references/nim/README.md for full endpoint docs and health checks.
    • External endpoints are opt-in only (explicit request or explicit URLs); only then skip the in-cluster deploy.
    • Never infer external mode from credential presence.
    • Never scale down/delete existing NIMs to free GPUs.
  4. Readiness guard

    bash
    osmo pool list --mode free
    osmo config show POD_TEMPLATE
    python3 scripts/pre_submit_guard.py --workflow assets/configs/osmo/<mode>.yaml
  5. Cache auto-remediation

    If pre_submit_guard.py reports cache failure, default action is to run:

    bash
    osmo workflow submit assets/configs/osmo/setup_model_cache.yaml \
      --set-string storage_url=<backend-prefix> path=data

    Then rerun pre_submit_guard.py and submit the target VDA flow only after it passes. Ask user only when backend/prefix is ambiguous or cache setup fails.

  6. Scheduling policy

    VDA templates schedule setup and workers on gpu_platform (no system pool dependency for user workloads).

Submit (all flows)

Every flow uses the same submit shape; only the workflow YAML changes. Choose the YAML for the requested flow, then run the command below. Full per-flow walkthroughs (stage matrix and flow details) live in the linked references.

FlowWorkflow YAMLWalkthrough
Augmentation + auto-labelingassets/configs/osmo/augmentation_and_al.yamlreferences/flows/augmentation_and_al.md
Auto-labeling onlyassets/configs/osmo/auto_labeling.yamlreferences/flows/auto_labeling.md
E2E (parallel)assets/configs/osmo/e2e.yamlreferences/flows/e2e.md
E2E (super-resolution gated)assets/configs/osmo/e2e_super_resolution.yamlreferences/flows/e2e_super_resolution.md
bash
SKILLS_DIR="$(cd "$(git rev-parse --show-toplevel)/skills/physical-ai-video-data-augmentation" && pwd)"
STAMP=$(cat /proc/sys/kernel/random/uuid | cut -c1-8)
osmo workflow submit assets/configs/osmo/<flow>.yaml \
  --pool <pool> \
  --set-string \
    dataset=<dataset> \
    run_id=run-$STAMP \
    storage_url=<backend-prefix> \
    gpu_platform=<gpu-platform> \
    video=<video-stem> \
    cosmos_model_cache_url=<backend-prefix>/data/models/cosmos_transfer \
    auto_labeling_model_cache_url=<backend-prefix>/data/models/auto_labeling \
    skills_dir="$SKILLS_DIR"

Compatibility note:

  • Use exactly one --set-string flag and pass all the key/value pairs after it.
  • Do not repeat --set/--set-string flags in the same command; some OSMO builds only honor the last occurrence.
  • Do not mix --set and --set-string in one submit command.
  • Pass explicit *_model_cache_url values to avoid nested-template interpolation differences across OSMO environments.
  • Do not brute-force permutations of flags. Use this shape directly.

Common optional overrides (append key/value pairs to the same --set-string list):

bash
cookbook=<scene_profile> \
vlm_url=<openai_base_url> \
llm_url=<openai_base_url> \
cosmos_model_cache_url=<url> \
auto_labeling_model_cache_url=<url>

The auto-labeling-only flow has no augmentation stage, so it omits cosmos_model_cache_url at runtime; passing it is harmless and keeps one submit shape across flows.

OSMO Monitoring

bash
# Workflow status + task states
osmo workflow query <workflow_id> --format-type json \
  | jq '{status, tasks: [.groups[].tasks[] | {name, status, exit_code}]}'

# Logs for a specific task
osmo workflow logs <workflow_id> --task <task_name> -n 200

# Output retrieval
osmo data list --no-pager <output_url>
osmo data download <output_url> <local_dir>/

For completion artifacts, always mirror the full run output into workspace:

bash
ROOT="$(git rev-parse --show-toplevel)"
RUN_LOCAL_DIR="$ROOT/media/vda/runs/<run_id>"
mkdir -p "$RUN_LOCAL_DIR"
osmo data download "<storage_url>/datasets/<dataset>-outputs/<run_id>/" "$RUN_LOCAL_DIR/"

For runs expected to exceed two minutes, send heartbeat updates at least every two minutes. For media evidence, emit one standalone MEDIA:<absolute-path> line per message bubble.

Execution continuity requirement:

  • Heartbeats must report progress while continuing work; they are status updates, not permission prompts.
  • Do not stop between green stages waiting for approval.
  • Pause only on blocking failures or explicit user stop/redirect.
  • If submit fails on interpolation, rerun once with the same canonical single-flag shape and corrected values; do not loop through ad-hoc flag experiments.

MEDIA formatting is strict:

  • Emit exactly one line: MEDIA:/absolute/path/to/file.mp4
  • Keep MEDIA: contiguous on a single line (never split across lines).
  • No extra text in the same bubble.
  • No code fences, bullets, or quotes around the directive.
  • If render fails: retry once from a stable workspace path, then emit PNG fallback.

Post-Run Comparison Evidence (required for augmented flows)

Applies to augmentation_and_al, e2e, and e2e_super_resolution after a successful run.

Required completion output (do not stop at raw output URLs):

  1. Stage full outputs + input video into workspace-local path:

    bash
    bash scripts/stage_run_artifacts.sh \
      --storage-url <storage_url> --dataset <dataset> --run-id <run_id> --video <video>
  2. Render side-by-side from that local run copy:

    bash
    bash scripts/render_side_by_side.sh \
      --run-local-dir "<repo>/media/vda/runs/<run_id>" --dataset <dataset> --video <video>
  3. Emit MEDIA from the local run copy and include:

    • augmentation summary from <run_local_dir>/setup_b0/configs/manifest.yaml (sampled_vars for <video>_aug0)
    • auto-labeling summary from <run_local_dir>/outputs/pseudo_labeled_augmented/<video>_aug0
    • for e2e / e2e_super_resolution, original-label summary from <run_local_dir>/outputs/pseudo_labeled/<video>

If ffmpeg is unavailable, emit input and augmented MEDIA from the same local run copy and still provide augmentation + auto-labeling summaries.

For demo runs (no user video provided), explicitly state that input came from nvidia/video-data-augmentation-demo.

Supporting files

Use these canonical locations:

  • Workflows: assets/configs/osmo/*.yaml
  • Runtime scripts: scripts/*.sh, scripts/*.py
  • Flow walkthroughs: references/flows/*.md
  • Setup and triage: references/setup.md, references/troubleshooting.md
  • Images and endpoint policy: references/container-images.md, references/nim/README.md
  • Cookbook tuning: assets/cookbooks/TUNING_GUIDE.md

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 102 other files (scripts, references, assets) in skills/physical-ai-video-data-augmentation of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • agents/openai.yaml
  • assets/configs/osmo/augmentation_and_al.yaml
  • assets/configs/osmo/augmentation_and_pl.yaml
  • assets/configs/osmo/auto_labeling.yaml
  • assets/configs/osmo/e2e.yaml
  • assets/configs/osmo/e2e_super_resolution.yaml
  • assets/configs/osmo/setup_model_cache.yaml
  • assets/cookbooks/FILE_INVENTORY.md
  • assets/cookbooks/TUNING_GUIDE.md
  • assets/cookbooks/city_traffic/README.md
  • assets/cookbooks/city_traffic/augmentation/augmentation.yaml
  • assets/cookbooks/city_traffic/augmentation/prompts
  • … and 89 more

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Physical AI Video Augmentation on OSMO next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Physical AI Video Augmentation on OSMO compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Physical AI Video Augmentation on OSMO this skillNVIDIA/skills3.5k—~4.7kAutomated safety check: NotesApache-2.0
Hugging Face Vision Trainerhuggingface/skills11k1 repos~7.5kAutomated safety check: PassApache-2.0
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Fla Triton To Gluonfla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMIT
Hugging Face LLM Trainerhuggingface/skills11k1 repos~7.2kAutomated safety check: PassApache-2.0

Similar skills

  • Hugging Face Vision Trainer

    huggingface/skills

    Official

    Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

    11k GitHub starsUsed in 1 repo~7.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hugging Face LLM Trainer

    huggingface/skills

    Official

    Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

    11k GitHub starsUsed in 1 repo~7.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs TAO Data Services gap analysis that compares ground-truth and predicted boxes to find weak images by per-class recall, precision and AP50.

    3.5k GitHub stars~1.8k tokensUpdated today
    Auto-check: notes

Questions about Physical AI Video Augmentation on OSMO

What does Physical AI Video Augmentation on OSMO do?

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. The agent picks one of four workflows from your intent (`auto_labeling`, `augmentation_and_al`, `e2e` or `e2e_super_resolution`), gives a rough execution-time overview, and runs preflight and readiness checks before anything is submitted. Submit-time values come from the active dataset backend, and it must never guess `storage_url`.

When should I use Physical AI Video Augmentation on OSMO?

Physical AI Video Augmentation on OSMO fits situations like: running a video data augmentation job on an OSMO cluster; auto-labeling a set of videos with pseudo labels; checking credentials and GPU pool readiness before submitting a workflow; downloading outputs and comparing augmented clips with the originals.

How do I install Physical AI Video Augmentation on OSMO in Claude Code?

Run `npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a claude-code`. Or copy the skill folder (skills/physical-ai-video-data-augmentation in NVIDIA/skills) into .claude/skills/physical-ai-video-data-augmentation in your project. Claude Code loads it when a task matches its description.

How do I install Physical AI Video Augmentation on OSMO in Codex?

Run `npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a codex`. Or copy the skill folder (skills/physical-ai-video-data-augmentation in NVIDIA/skills) into .agents/skills/physical-ai-video-data-augmentation in your project. Codex loads it when a task matches its description.

Can I use Physical AI Video Augmentation on OSMO in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/physical-ai-video-data-augmentation, .gemini/skills/physical-ai-video-data-augmentation, .github/skills/physical-ai-video-data-augmentation and .opencode/skills/physical-ai-video-data-augmentation in your project.

What does Physical AI Video Augmentation on OSMO need to run?

Going by SKILL.md and its folder, Physical AI Video Augmentation on OSMO needs the command-line tools its instructions call (bash, git, python3 and jq) and credentials named NGC_API_KEY, NGC_CLI_API_KEY, NVIDIA_API_KEY and OPENAI_API_KEY. Our summary lists: The `osmo` CLI, logged in with a default profile and data credential profile; A Hugging Face token for gated model weights; At least one online GPU pool on OSMO.

Does Physical AI Video Augmentation on OSMO access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Physical AI Video Augmentation on OSMO safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Physical AI Video Augmentation on OSMO use?

Physical AI Video Augmentation on OSMO is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Physical AI Video Augmentation on OSMO use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.5k tokens, read only when the agent opens those files.

What are the alternatives to Physical AI Video Augmentation on OSMO?

Skills that share tags, products or a category with Physical AI Video Augmentation on OSMO: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Physical AI Video Augmentation on OSMO?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.