LLM Torch Profiler Analysis
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures.
$ npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-train-depth-anything-v2 --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-train-depth-anything-v2 .claude/skills/tao-train-depth-anything-v2 && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-train-depth-anything-v2" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-depth-anything-v2 into .claude/skills/tao-train-depth-anything-v2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-depth-anything-v2", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-train-depth-anything-v2Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-train-depth-anything-v2 --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-train-depth-anything-v2 .agents/skills/tao-train-depth-anything-v2 && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-train-depth-anything-v2" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-depth-anything-v2 into .agents/skills/tao-train-depth-anything-v2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-depth-anything-v2", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-train-depth-anything-v2 --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-train-depth-anything-v2 .cursor/skills/tao-train-depth-anything-v2 && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-train-depth-anything-v2" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-depth-anything-v2 into .cursor/skills/tao-train-depth-anything-v2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-depth-anything-v2", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-train-depth-anything-v2--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-train-depth-anything-v2 --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-train-depth-anything-v2 .gemini/skills/tao-train-depth-anything-v2 && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-train-depth-anything-v2" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-depth-anything-v2 into .gemini/skills/tao-train-depth-anything-v2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-depth-anything-v2", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-train-depth-anything-v2Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-train-depth-anything-v2 .github/skills/tao-train-depth-anything-v2 && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-depth-anything-v2" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-depth-anything-v2 into .github/skills/tao-train-depth-anything-v2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-depth-anything-v2", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-train-depth-anything-v2 --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-train-depth-anything-v2 .opencode/skills/tao-train-depth-anything-v2 && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-train-depth-anything-v2" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-train-depth-anything-v2 into .opencode/skills/tao-train-depth-anything-v2/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-train-depth-anything-v2", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-train-depth-anything-v2Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures.
Tao Train Depth Anything V2 is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures. Predicts per-pixel depth from single RGB images. Use when training, evaluating, exporting, or running inference for a TAO monocular depth model. Trigger phrases include "train monocular depth", "DepthAnything v2", "metric depth from single image", "monocular depth estimation".
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 31 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires docker + nvidia-container-toolkit.
It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
dockerFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires docker + nvidia-container-toolkit.
From compatibility in the SKILL.md frontmatter.
Tao Train Depth Anything V2 loads about 3.9k tokens when it runs, and up to ~22k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 1,556 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 1,556 words, ~3,906 tokens.
.claude/skills/tao-train-depth-anything-v2/SKILL.md (or your agent's skills folder). This skill also uses 28 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures. Predicts per-pixel depth from single RGB images.
Pretrained checkpoint loading varies by model variant and use case — see the Pretrained checkpoint loading — use case matrix in references/parameters.md.
The mono and stereo skills both invoke the unified TAO depth_net CLI inside the container; the mono/stereo family is selected via model.model_type (see references/parameters.md).
For TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and TensorRT inference), read references/tao-deploy-depth-anything-v2.md first. The deploy spec template lives in this skill's references/spec_template_deploy.yaml.
PyT actions packaged by this model skill: train, evaluate, inference, export, and quantize. The PyT depth_net entrypoint does not accept a PyT-side gen_trt_engine action in the current TAO image. The gen_trt_engine action metadata must run with the TAO Deploy container, and the deploy workflow remains the deploy-specific entrypoint.
This model is AutoML-enabled at the model layer. Before handling any train-stage request, read references/skill_info.yaml and resolve the run override from either an explicit automl_policy value or the user's workflow request. Use automl_policy: on by default and only expose on / off in new launch prompts. Treat phrases like "turn off AutoML", "disable AutoML", "no HPO", or "plain training" as automl_policy: off for this run only. When automl_policy: on, automl_enabled: true, and both schemas/train.schema.json and references/spec_template_train.yaml are packaged, route the train action through tao-skill-bank:tao-run-automl by default with this model's skill_dir. Preserve workflow/application overrides for datasets, specs, output directories, GPU/platform settings, parent checkpoints, and automl_policy. Use direct model training only when automl_policy: off or the packaged train schema/template is missing; in the missing-schema case, report that AutoML is enabled but not runnable for this model until schemas are generated.
Non-train actions such as evaluate, inference, export, and deploy flows stay in this model skill. The per-run automl_policy override does not change model metadata.
Your dataset (RGB images + GT depth files) must be reachable from inside the container:
S3_TRAIN / S3_EVAL placeholders shown in Typical Spec Overrides). The runner handles S3 → container-path mounting transparently.docker run (e.g. local testing): mount the host dataset root read-only at the same in-container path:docker run ... -v <host_data_root>:<host_data_root>:ro <container> ...The same accessibility requirement applies to the <output_dir> written by all actions.
Per-line annotation file referenced by data_sources[*].data_file:
| Columns | Format | Use |
|---|---|---|
| 1 | <image> | Mono inference (no GT) |
| 2 | <image> <gt_depth> | Mono with GT |
Do not pass stereo annotation rows such as <left_image> <right_image> <gt_depth> directly to mono train/evaluate/inference. If only a stereo depth
dataset is available, derive a mono annotation file by keeping the left image
and GT depth columns, then mount or stage the image/depth archive at the same
container paths referenced by that derived annotation file.
If you already have one, point to it. Otherwise generate via depth_net convert:
depth_net convert -e <convert_spec.yaml>convert_spec.yaml template:
results_dir: <directory where generated annotation files are written>
data_root: <directory whose immediate children are scene/sample folders that contain your image+depth files; convert walks data_root recursively but expects per-scene subdirectories at one level below>
image_dir_pattern: [<substring matching left/RGB image paths>]
depth_dir_pattern: [<substring matching GT depth paths>]
image_extension: '' # optional .endswith filter, e.g. '.jpg'
depth_extension: '' # optional, swapped during depth derivation, e.g. '.png'
split_ratio: 0.0 # 0.0/1.0 = test-only; 0.8 = 80/20 train+valconvert walks data_root recursively, selects paths whose path-string contains all substrings in image_dir_pattern (AND-filter), then derives the depth path by replacing image_dir_pattern[0] with depth_dir_pattern[0] and image_extension with depth_extension. Inspect your dataset's directory layout and identify the substring distinguishing RGB images from depth files (e.g. rgb_ vs sync_depth_).
data_root must point at the parent that contains the per-scene subdirectories (e.g. for NYU eval, use /data/nyu_v2/eval/test, not /data/nyu_v2/eval/test/bathroom — the latter limits the walk to a single scene). Always include the leading dot in image_extension / depth_extension (e.g. '.jpg' not 'jpg'); the substring swap is form-sensitive and a mismatch silently corrupts derived paths.
model_type and dataset_name based on your dataDefault — generic class for each task:
| Data category | model_type | dataset_name |
|---|---|---|
| Disparity-encoded data (pixels) | RelativeDepthAnything | RelativeMonoDataset |
| Metric depth (meters) | MetricDepthAnything | MetricMonoDataset |
| Mono inference (no GT, any image) | matches train choice | RelativeMonoDataset or MetricMonoDataset |
Dataset-specific class — switch when the data needs preprocessing the generic class does not perform:
| Special case | model_type | dataset_name | What the class adds |
|---|---|---|---|
NYU sync_depth_*.png (raw uint16 millimetres) — relative | RelativeDepthAnything | NYUDV2Relative | mm→m unit conversion + Eigen evaluation crop |
NYU sync_depth_*.png (raw uint16 millimetres) — metric | MetricDepthAnything | NYUDV2 | same |
Using a generic class on data that requires unit conversion (e.g. raw NYU uint16 PNGs) results in an empty valid mask and silent train_loss = NaN. Match the class to your data's encoding.
For relative mono data (RelativeMonoDataset or NYUDV2Relative), leave dataset.min_depth and dataset.max_depth unset or set both to null. Non-null metric depth ranges are passed into the relative dataset constructor and fail with BaseRelativeMonoDataset.__init__() got an unexpected keyword argument 'min_depth'.
Copy the action block from Typical Spec Overrides (references/spec-overrides.md). Replace:
model.model_type from Step 2dataset.<...>.data_sources[*].dataset_name from Step 2data_sources[*].data_file with the path from Step 1 (S3 path under SDK runner, host path for direct docker)references/finetuning-recipes.md.For mono training set train.precision: fp32 (recommended) or bf16 (Ampere SM80+, alternative).
Create writable home/cache directories inside the mounted output path before using
--user. Some TAO containers do not have an /etc/passwd entry for the host UID,
and PyTorch / matplotlib need writable cache paths when running as that UID.
mkdir -p <output_dir>/home \
<output_dir>/.cache/matplotlib \
<output_dir>/.cache/torchinductor \
<output_dir>/.cache/xdgdocker run --gpus 'device=0' --shm-size 16G --shm-size=8g \
--user "$(id -u):$(id -g)" \
-e USER="$(id -un)" \
-e LOGNAME="$(id -un)" \
-e HOME=<output_dir>/home \
-e MPLCONFIGDIR=<output_dir>/.cache/matplotlib \
-e TORCHINDUCTOR_CACHE_DIR=<output_dir>/.cache/torchinductor \
-e XDG_CACHE_HOME=<output_dir>/.cache/xdg \
-v <data_root>:<data_root>:ro \
-v <output_dir>:<output_dir> \
<container> \
depth_net <action> -e <spec.yaml>Without --user "$(id -u):$(id -g)" the container writes outputs as nobody:nogroup, blocking host-side cleanup and retry.
status.json kpi block populatedtrain: inspect per-step train_loss directly — the entrypoint reports Execution status: PASS even when train_loss = NaN (see the Metric Variant Finetuning Recipe → Sanity-run PASS criteria in references/finetuning-recipes.md)evaluate / inference: artifacts under results_dirFor TAO Deploy TensorRT actions (gen_trt_engine, TensorRT evaluate, and TensorRT inference), read references/tao-deploy-depth-anything-v2.md first. Deploy spec templates live in this skill's references/ folder with the spec_template_deploy_*.yaml prefix.
dataset_name values for mono data_sources (case-insensitive): ThreeDVLM, FSD, NvCLIP, IssacStereo, Crestereo, Middlebury, NYUDV2, NYUDV2Relative, RelativeMonoDataset, MetricMonoDataset. NYUDV2 carries metric depth GT (meters) — pair with MetricDepthAnything; NYUDV2Relative is the same data with relative-depth conventions — pair with RelativeDepthAnything.val/d1 (maximize), val/loss (minimize).val/d1 as the primary monitor and maximize
it. Mono d1 is Delta-1
accuracy: the fraction of valid pixels where
max(pred/target, target/pred) < 1.25. Do not apply the stereo d1 error-rate
direction to this mono metric. val/loss can be emitted as NaN even when
the trainer exits successfully and writes a usable checkpoint, so it is not
a reliable AutoML objective unless the run's status metrics show a finite
value.train.optim.lr and train.optim.weight_decay for the default relative-depth train workflow. Do not search dataset.val_dataset, dataset.test_dataset, or dataset.infer_dataset augmentation fields because they change scoring/non-train behavior rather than training. Do not search model.corr_radius, model.cv_group, or model.volume_dim for mono; those fields belong to the stereo architecture and are inert for RelativeDepthAnything.| Action | Spec Key | Source | Files | List? |
|---|---|---|---|---|
| evaluate | dataset.test_dataset.data_sources | eval_dataset | data_file: annotations.txt + dataset_name | Yes |
| inference | dataset.infer_dataset.data_sources | inference_dataset | data_file: annotations.txt + dataset_name | Yes |
| quantize | dataset.train_dataset.data_sources | train_datasets | data_file: annotations.txt + dataset_name | Yes |
| quantize | dataset.val_dataset.data_sources | eval_dataset | data_file: annotations.txt + dataset_name | Yes |
| quantize | dataset.quant_calibration_dataset.images_dir | train_datasets | images.tar.gz | No |
| train | dataset.train_dataset.data_sources | train_datasets | data_file: annotations.txt + dataset_name | Yes |
| train | dataset.val_dataset.data_sources | eval_dataset | data_file: annotations.txt + dataset_name | Yes |
Data source overrides are mandatory for every action — construct data source paths from the Per-Action Dataset Requirements table above and include them in spec_overrides. Each data_sources entry is a dict with two mandatory fields: data_file and dataset_name. See references/spec-overrides.md for the full per-action override blocks (train, evaluate, export, inference, quantize), the S3_TRAIN / S3_EVAL placeholders, the relative-variant precision recommendation, and the quantize known-issue note.
Optional. Val dataset configured via dataset.val_dataset.data_sources (each entry needs data_file and dataset_name).
See references/parameters.md for the full parameter glossary (model, train, dataset, export, and inference keys with options, defaults, and sources) and the Pretrained checkpoint loading — use case matrix.
See references/finetuning-recipes.md for:
RelativeDepthAnything checkpoint (lr 5e-6, LambdaLR, sanity-vs-convergent guidance, deploy LSQ alignment note).normalize_depth/min_depth/max_depth) required in train AND export specs, trainer-enforced defaults, precision, the 1-epoch sanity-run override, and the Sanity-run PASS criteria with the NaN-mitigation order.Launch method: Lightning-managed (single python process, Lightning spawns workers).
| Spec Key | Description | Default |
|---|---|---|
train.num_gpus | Number of GPUs | 1 |
train.gpu_ids | GPU device indices | [0] |
train.num_nodes | Number of nodes | 1 |
train.distributed_strategy | ddp or fsdp | ddp |
ddp with activation checkpointing: find_unused_parameters=Falseddp without: find_unused_parameters=Truefsdp forces precision to FP16Multi-node env vars (set by orchestrator): WORLD_SIZE, NODE_RANK, MASTER_ADDR, MASTER_PORT, NUM_GPU_PER_NODE.
fp32. BF16 is supported on Ampere SM80+ hardware, but keep smoke tests on FP32 unless the user explicitly requests BF16.Minimum 1 GPU(s), recommended 2 GPU(s). 24GB+ VRAM per GPU. ViT-Large encoder is memory intensive. Use fp32 (recommended) or bf16 (Ampere SM80+, alternative) for training. Activation checkpointing is available for larger inputs.
See references/troubleshooting.md for the full error-pattern catalog (depth range mismatch, relative dataset rejecting min_depth, missing pretrained weights, encoder key location, dataset_name not in struct, depth_net_mono not found, metric variant hyperparameter sourcing, and export ONNX overwrite).
See references/spec-param-inference.md for the model-specific inference mappings (the TAO Core depth_net_mono.config.json action table), checkpoint-file naming under <results_dir>/train/, the dn_model_latest.pth policy, the parent-gen_trt_engine rationale, and the parent_model / parent_job_id resolution rules.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 28 other files (references) in skills/tao-train-depth-anything-v2 of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Tao Train Depth Anything V2 next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Train Depth Anything V2 this skillNVIDIA/skills | 3.6k | — | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| LLM Torch Profiler Analysissgl-project/sglang | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| Skill InspectorNVIDIA/SkillSpector | 20k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Megatron-LM Container and Dependency SetupNVIDIA/Megatron-LM | 18k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | |
| Embeddings via 9Routerdecolua/9router | 31k | — | ~604 | Automated safety check: Pass | MIT | |
| Megatron-LM Base Image BumpNVIDIA/Megatron-LM | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 |
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
NVIDIA/SkillSpector
Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.
NVIDIA/Megatron-LM
Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.
decolua/9router
Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
NVIDIA/NemoClaw
Remove bracketed NemoClaw tags from GitHub issue and PR titles.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures. Tao Train Depth Anything V2 is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Monocular depth estimation using Metric Depth Anything v2 or Relative Depth Anything architectures.
Tao Train Depth Anything V2 fits situations like: running inference for a TAO monocular depth model; phrases include train monocular depth; depthAnything v2; metric depth from single image.
Run `npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a claude-code`. Or copy the skill folder (skills/tao-train-depth-anything-v2 in NVIDIA/skills) into .claude/skills/tao-train-depth-anything-v2 in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a codex`. Or copy the skill folder (skills/tao-train-depth-anything-v2 in NVIDIA/skills) into .agents/skills/tao-train-depth-anything-v2 in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-train-depth-anything-v2 -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-train-depth-anything-v2, .gemini/skills/tao-train-depth-anything-v2, .github/skills/tao-train-depth-anything-v2 and .opencode/skills/tao-train-depth-anything-v2 in your project.
Going by SKILL.md and its folder, Tao Train Depth Anything V2 needs the command-line tools its instructions call (docker). Our summary lists: Docker. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires docker + nvidia-container-toolkit..
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Train Depth Anything V2 is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Train Depth Anything V2: LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Skill Inspector (NVIDIA/SkillSpector, 20k stars), Megatron-LM Container and Dependency Setup (NVIDIA/Megatron-LM, 18k stars) and Embeddings via 9Router (decolua/9router, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.