Hugging Face Vision Trainer
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
$ npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills physical-ai-video-data-augmentation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/physical-ai-video-data-augmentation .claude/skills/physical-ai-video-data-augmentation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "physical-ai-video-data-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physical-ai-video-data-augmentation into .claude/skills/physical-ai-video-data-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physical-ai-video-data-augmentation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/physical-ai-video-data-augmentationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills physical-ai-video-data-augmentation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/physical-ai-video-data-augmentation .agents/skills/physical-ai-video-data-augmentation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "physical-ai-video-data-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physical-ai-video-data-augmentation into .agents/skills/physical-ai-video-data-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physical-ai-video-data-augmentation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills physical-ai-video-data-augmentation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/physical-ai-video-data-augmentation .cursor/skills/physical-ai-video-data-augmentation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "physical-ai-video-data-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physical-ai-video-data-augmentation into .cursor/skills/physical-ai-video-data-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physical-ai-video-data-augmentation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/physical-ai-video-data-augmentation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills physical-ai-video-data-augmentation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/physical-ai-video-data-augmentation .gemini/skills/physical-ai-video-data-augmentation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "physical-ai-video-data-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physical-ai-video-data-augmentation into .gemini/skills/physical-ai-video-data-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physical-ai-video-data-augmentation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills physical-ai-video-data-augmentationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/physical-ai-video-data-augmentation .github/skills/physical-ai-video-data-augmentation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "physical-ai-video-data-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physical-ai-video-data-augmentation into .github/skills/physical-ai-video-data-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physical-ai-video-data-augmentation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills physical-ai-video-data-augmentation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/physical-ai-video-data-augmentation .opencode/skills/physical-ai-video-data-augmentation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "physical-ai-video-data-augmentation" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physical-ai-video-data-augmentation into .opencode/skills/physical-ai-video-data-augmentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physical-ai-video-data-augmentation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
physical-ai-video-data-augmentationOrchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
The agent picks one of four workflows from your intent (`auto_labeling`, `augmentation_and_al`, `e2e` or `e2e_super_resolution`), gives a rough execution-time overview, and runs preflight and readiness checks before anything is submitted. Submit-time values come from the active dataset backend, and it must never guess `storage_url`. It then submits with explicit interpolation values, monitors to completion, retrieves outputs and provides side-by-side comparison evidence for augmented flows.
Prerequisites are checked first: an optional NGC API key, a Hugging Face token for gated Cosmos and SeedVR weights, the `osmo` CLI logged in with a default profile and a matching data credential profile, and at least one online GPU pool. Missing secrets surface as `USER_INPUT_REQUIRED` from `scripts/preflight_credentials.sh`. The folder ships 109 files, including OSMO configs and cookbooks. Questions about tuning inside the container are out of scope.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
bashgitpython3jqFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NGC_API_KEYNGC_CLI_API_KEYNVIDIA_API_KEYOPENAI_API_KEYVLM_API_KEYLLM_API_KEYHF_TOKENHUGGING_FACE_HUB_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Physical AI Video Augmentation on OSMO loads about 4.7k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 1,809 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
light does not require a workload-local `.env`. Runtime interpolation isAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,809 words, ~4,747 tokens.
.claude/skills/physical-ai-video-data-augmentation/SKILL.md (or your agent's skills folder). This skill also uses 102 other files; get the full folder from GitHub.Default workflow skill for VDA execution on OSMO. It owns flow selection, preflight, cache readiness, inference-path decisions, submit-time interpolation, monitoring, and output retrieval. Component skills are consult-only.
Run the end-to-end VDA workflow safely and reproducibly from preflight to output download.
Do NOT use this skill for container-internal tuning-only questions.
Confirm these before running preflight or any submit. Missing required secrets
surface as USER_INPUT_REQUIRED: from scripts/preflight_credentials.sh.
| Requirement | How it is satisfied | Used for |
|---|---|---|
| NGC API key (optional) | NGC_API_KEY, NGC_CLI_API_KEY, or compatible nvapi-* token in NVIDIA_API_KEY/OPENAI_API_KEY/VLM_API_KEY/LLM_API_KEY | Optional for nvcr_io credential refresh and NGC REST scope probe; default VDA image refs are validated via workflow registry probes |
| Hugging Face token | HF_TOKEN (or HUGGING_FACE_HUB_TOKEN), or a cached token at ~/.cache/huggingface/token | Creates the OSMO hf_token credential; pulls gated Cosmos/SeedVR weights |
| OSMO CLI access | osmo on PATH, logged in, with a default profile and a registered DATA credential profile matching storage_url | Submitting/monitoring workflows and listing/downloading objects |
| GPU pool | At least one ONLINE pool in osmo pool list --mode free; POD_TEMPLATE carries GPU toleration/selectors | Scheduling setup + worker tasks |
Optional (only for the strict NGC org/team probe): NGC_ORG + NGC_TEAM
(or NGC_CLI_ORG / NGC_CLI_TEAM). External VLM/LLM endpoint keys are validated
separately, not by preflight.
Key handling rule: nvapi-* tokens are first-class inputs for nvcr_io.
Never reject by token prefix alone; use workflow registry probe results as
source of truth.
auto_labeling, augmentation_and_al, e2e,
e2e_super_resolution) from user intent.storage_url).Use run_script(...) for script execution. Canonical examples:
run_script("bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/augmentation_and_al.yaml")
run_script("python3 scripts/pre_submit_guard.py --workflow assets/configs/osmo/auto_labeling.yaml")
run_script("bash scripts/prepare_demo_assets.sh /srv/sdg/data/vda_inputs")Use script-level --help for exact arguments.
| Script | Role |
|---|---|
scripts/preflight_credentials.sh | Secrets/control-plane preflight and workflow image access checks |
scripts/pre_submit_guard.py | Submit-time interpolation, cache, and dataset safety checks |
scripts/prepare_demo_assets.sh | Demo video pull + flatten for default demo path |
scripts/generate_configs.py | Setup-time config and cookbook projection generation |
scripts/cosmos_worker.sh | Augmentation worker execution |
scripts/pl_original_worker.sh | Original-video auto-labeling worker execution |
scripts/pl_augmented_worker.sh | Augmented-video auto-labeling worker execution |
scripts/osmo_barrier.py | Multi-node barrier synchronization |
scripts/stage_run_artifacts.sh | Local mirror of full run output + input video |
scripts/render_side_by_side.sh | Side-by-side comparison render from local artifacts |
| Flow | OSMO YAML | Group sequence | Typical use |
|---|---|---|---|
augmentation_and_al | assets/configs/osmo/augmentation_and_al.yaml | setup -> augmentation -> auto_labeling_augmented | Augment one or more videos, then auto-label augmented outputs |
auto_labeling | assets/configs/osmo/auto_labeling.yaml | setup -> auto_labeling | Label original videos only |
e2e | assets/configs/osmo/e2e.yaml | setup -> (auto_labeling_original + augmentation) -> auto_labeling_augmented | Throughput-first path |
e2e_super_resolution | assets/configs/osmo/e2e_super_resolution.yaml | setup -> auto_labeling_original -> augmentation -> auto_labeling_augmented | Sequential path with SR gate before augmentation |
Legacy alias assets/configs/osmo/augmentation_and_pl.yaml remains for
backwards compatibility.
| User intent | Workflow |
|---|---|
| "Label my source videos" / "PL-only" / "no augmentation" | auto_labeling |
| "Create augmented videos and label them" | augmentation_and_al |
| "Run the full pipeline quickly" | e2e |
| "Run full pipeline, but gate on SR-enhanced originals first" | e2e_super_resolution |
Default to autonomy: ask only when missing information blocks execution.
scripts/prepare_demo_assets.sh)
and continue with dataset=vda-demo.augmentation_and_al.setup_model_cache.yaml, rerun pre-submit guard, and
continue automatically on success.| Missing input | Why it matters | Ask |
|---|---|---|
USER_INPUT_REQUIRED from preflight | Required secret is missing | Ask one concise unblock question for exactly the missing value(s) |
| Storage backend prefix cannot be derived from the active dataset/upload root | Wrong scheme causes runtime storage auth mismatch | "What is the backend-native root prefix for this run?" |
| No ONLINE GPU pool/platform can be selected | Workflow cannot schedule setup/workers | "Which GPU pool/platform should this run target?" |
scripts/prepare_demo_assets.sh (HF dataset flow) without asking extra
source-selection questions.nvidia/video-data-augmentation-demo) for the
default demo path.Collect only missing values:
dataset_url or local upload
folder; otherwise default to VDA demo assets and proceed).auto_labeling, augmentation_and_al, e2e, e2e_super_resolution);
default to augmentation_and_al when unspecified.gpu_platform for all VDA resources (auto-select an ONLINE platform
when unambiguous; ask only when no valid option exists).Do not guess gpu_platform (for example microk8s). Use the exact current
platform label shown by osmo pool list --mode free (for example gpu).
Generate run stamp before each submit:
STAMP=$(cat /proc/sys/kernel/random/uuid | cut -c1-8)
RUN_ID="run-$STAMP"Before running any mutating command (osmo credential set, NIM install/repair,
cache workflow submit, or target VDA workflow submit), provide a short ETA
overview to the user.
Keep it concise (one short paragraph or 4-6 bullets) and include:
Baseline ranges (from observed MicroK8s + OSMO runs):
| Phase | Typical duration |
|---|---|
| Credentials + preflight | ~1-2 min |
| NIM deploy/download/warmup (if needed) | ~10-15 min |
| Demo assets download/upload (if demo path) | ~1-3 min |
| Model cache population (if needed) | ~15-25 min |
| Workflow submit + queue/start | ~1-3 min |
Workflow runtime ranges after submit:
| Flow | Typical runtime |
|---|---|
auto_labeling | ~6-15 min |
augmentation_and_al | ~20-35 min |
e2e | ~22-40 min |
e2e_super_resolution | ~25-45 min |
Cold-start end-to-end runs are commonly ~45-80 min; warm-start runs are usually ~20-45 min depending on flow and video length.
Credential and control-plane preflight
bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/<mode>.yamlRestricted egress:
bash scripts/preflight_credentials.sh --no-probe --workflow assets/configs/osmo/<mode>.yamlPreflight does not require a workload-local .env. Runtime interpolation is
driven by submit-time values (dataset, run_id, gpu_platform, video,
storage_url, skills_dir) supplied in one --set-string list.
Passing --workflow validates pull access for the active workflow image refs
(workflow.groups[].tasks[].image) using anonymous bearer access with
credential fallback when provided.
If replacement NGC/HF secrets are provided in env, preflight refreshes
existing nvcr_io / hf_token automatically when present. Use --refresh to force
overwrite even when no new env secrets were supplied:
bash scripts/preflight_credentials.sh --workflow assets/configs/osmo/<mode>.yaml --refreshIf output contains USER_INPUT_REQUIRED:, ask one concise unblock question
and stop.
On workflow image 401/403, report registry access failure after probe
checks on the listed image refs; do not claim a key family (for example
nvapi-*) is categorically unsupported.
Storage interpolation policy
storage_url must be derived from the actual dataset/upload backend for the
current run.
dataset_url=azure://storiondevxah69/osmo-workflows/datasets/vda-demo
storage_url=azure://storiondevxah69/osmo-workflows
dataset=vda-demoNever silently default to stale s3:// values on non-S3 backends.
Inference policy (non-negotiable)
export NIM_SERVICES="qwen3-vl qwen25-14b"
skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/inference-nim-operator/scripts/install.shreferences/nim/README.md for full endpoint docs and health checks.Readiness guard
osmo pool list --mode free
osmo config show POD_TEMPLATE
python3 scripts/pre_submit_guard.py --workflow assets/configs/osmo/<mode>.yamlCache auto-remediation
If pre_submit_guard.py reports cache failure, default action is to run:
osmo workflow submit assets/configs/osmo/setup_model_cache.yaml \
--set-string storage_url=<backend-prefix> path=dataThen rerun pre_submit_guard.py and submit the target VDA flow only after it
passes. Ask user only when backend/prefix is ambiguous or cache setup fails.
Scheduling policy
VDA templates schedule setup and workers on gpu_platform (no system pool
dependency for user workloads).
Every flow uses the same submit shape; only the workflow YAML changes. Choose the YAML for the requested flow, then run the command below. Full per-flow walkthroughs (stage matrix and flow details) live in the linked references.
| Flow | Workflow YAML | Walkthrough |
|---|---|---|
| Augmentation + auto-labeling | assets/configs/osmo/augmentation_and_al.yaml | references/flows/augmentation_and_al.md |
| Auto-labeling only | assets/configs/osmo/auto_labeling.yaml | references/flows/auto_labeling.md |
| E2E (parallel) | assets/configs/osmo/e2e.yaml | references/flows/e2e.md |
| E2E (super-resolution gated) | assets/configs/osmo/e2e_super_resolution.yaml | references/flows/e2e_super_resolution.md |
SKILLS_DIR="$(cd "$(git rev-parse --show-toplevel)/skills/physical-ai-video-data-augmentation" && pwd)"
STAMP=$(cat /proc/sys/kernel/random/uuid | cut -c1-8)
osmo workflow submit assets/configs/osmo/<flow>.yaml \
--pool <pool> \
--set-string \
dataset=<dataset> \
run_id=run-$STAMP \
storage_url=<backend-prefix> \
gpu_platform=<gpu-platform> \
video=<video-stem> \
cosmos_model_cache_url=<backend-prefix>/data/models/cosmos_transfer \
auto_labeling_model_cache_url=<backend-prefix>/data/models/auto_labeling \
skills_dir="$SKILLS_DIR"Compatibility note:
--set-string flag and pass all the key/value pairs after it.--set/--set-string flags in the same command; some OSMO builds
only honor the last occurrence.--set and --set-string in one submit command.*_model_cache_url values to avoid nested-template interpolation
differences across OSMO environments.Common optional overrides (append key/value pairs to the same --set-string list):
cookbook=<scene_profile> \
vlm_url=<openai_base_url> \
llm_url=<openai_base_url> \
cosmos_model_cache_url=<url> \
auto_labeling_model_cache_url=<url>The auto-labeling-only flow has no augmentation stage, so it omits
cosmos_model_cache_url at runtime; passing it is harmless and keeps one submit
shape across flows.
# Workflow status + task states
osmo workflow query <workflow_id> --format-type json \
| jq '{status, tasks: [.groups[].tasks[] | {name, status, exit_code}]}'
# Logs for a specific task
osmo workflow logs <workflow_id> --task <task_name> -n 200
# Output retrieval
osmo data list --no-pager <output_url>
osmo data download <output_url> <local_dir>/For completion artifacts, always mirror the full run output into workspace:
ROOT="$(git rev-parse --show-toplevel)"
RUN_LOCAL_DIR="$ROOT/media/vda/runs/<run_id>"
mkdir -p "$RUN_LOCAL_DIR"
osmo data download "<storage_url>/datasets/<dataset>-outputs/<run_id>/" "$RUN_LOCAL_DIR/"For runs expected to exceed two minutes, send heartbeat updates at least every
two minutes. For media evidence, emit one standalone MEDIA:<absolute-path>
line per message bubble.
Execution continuity requirement:
MEDIA formatting is strict:
MEDIA:/absolute/path/to/file.mp4MEDIA: contiguous on a single line (never split across lines).Applies to augmentation_and_al, e2e, and e2e_super_resolution after a
successful run.
Required completion output (do not stop at raw output URLs):
Stage full outputs + input video into workspace-local path:
bash scripts/stage_run_artifacts.sh \
--storage-url <storage_url> --dataset <dataset> --run-id <run_id> --video <video>Render side-by-side from that local run copy:
bash scripts/render_side_by_side.sh \
--run-local-dir "<repo>/media/vda/runs/<run_id>" --dataset <dataset> --video <video>Emit MEDIA from the local run copy and include:
<run_local_dir>/setup_b0/configs/manifest.yaml
(sampled_vars for <video>_aug0)<run_local_dir>/outputs/pseudo_labeled_augmented/<video>_aug0e2e / e2e_super_resolution, original-label summary from
<run_local_dir>/outputs/pseudo_labeled/<video>If ffmpeg is unavailable, emit input and augmented MEDIA from the same local
run copy and still provide augmentation + auto-labeling summaries.
For demo runs (no user video provided), explicitly state that input came from
nvidia/video-data-augmentation-demo.
Use these canonical locations:
assets/configs/osmo/*.yamlscripts/*.sh, scripts/*.pyreferences/flows/*.mdreferences/setup.md, references/troubleshooting.mdreferences/container-images.md, references/nim/README.mdassets/cookbooks/TUNING_GUIDE.md© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 102 other files (scripts, references, assets) in skills/physical-ai-video-data-augmentation of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
Physical AI Video Augmentation on OSMO next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Physical AI Video Augmentation on OSMO this skillNVIDIA/skills | 3.5k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | |
| Hugging Face Vision Trainerhuggingface/skills | 11k | 1 repos | ~7.5k | Automated safety check: Pass | Apache-2.0 | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3.3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Fla Triton To Gluonfla-org/flash-linear-attention | 5.8k | — | ~4.2k | Automated safety check: Pass | MIT | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 1 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
fla-org/flash-linear-attention
Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
NVIDIA/skills
Runs TAO Data Services gap analysis that compares ground-truth and predicted boxes to find weak images by per-class recall, precision and AP50.
Works with
Categories
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. The agent picks one of four workflows from your intent (`auto_labeling`, `augmentation_and_al`, `e2e` or `e2e_super_resolution`), gives a rough execution-time overview, and runs preflight and readiness checks before anything is submitted. Submit-time values come from the active dataset backend, and it must never guess `storage_url`.
Physical AI Video Augmentation on OSMO fits situations like: running a video data augmentation job on an OSMO cluster; auto-labeling a set of videos with pseudo labels; checking credentials and GPU pool readiness before submitting a workflow; downloading outputs and comparing augmented clips with the originals.
Run `npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a claude-code`. Or copy the skill folder (skills/physical-ai-video-data-augmentation in NVIDIA/skills) into .claude/skills/physical-ai-video-data-augmentation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a codex`. Or copy the skill folder (skills/physical-ai-video-data-augmentation in NVIDIA/skills) into .agents/skills/physical-ai-video-data-augmentation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill physical-ai-video-data-augmentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/physical-ai-video-data-augmentation, .gemini/skills/physical-ai-video-data-augmentation, .github/skills/physical-ai-video-data-augmentation and .opencode/skills/physical-ai-video-data-augmentation in your project.
Going by SKILL.md and its folder, Physical AI Video Augmentation on OSMO needs the command-line tools its instructions call (bash, git, python3 and jq) and credentials named NGC_API_KEY, NGC_CLI_API_KEY, NVIDIA_API_KEY and OPENAI_API_KEY. Our summary lists: The `osmo` CLI, logged in with a default profile and data credential profile; A Hugging Face token for gated model weights; At least one online GPU pool on OSMO.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Physical AI Video Augmentation on OSMO is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.5k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Physical AI Video Augmentation on OSMO: Hugging Face Vision Trainer (huggingface/skills, 11k stars), Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.