GPU Optimizer
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and…
$ npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills physicsnemo-shard-tensor --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/physicsnemo-shard-tensor .claude/skills/physicsnemo-shard-tensor && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "physicsnemo-shard-tensor" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physicsnemo-shard-tensor into .claude/skills/physicsnemo-shard-tensor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physicsnemo-shard-tensor", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/physicsnemo-shard-tensorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills physicsnemo-shard-tensor --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/physicsnemo-shard-tensor .agents/skills/physicsnemo-shard-tensor && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "physicsnemo-shard-tensor" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physicsnemo-shard-tensor into .agents/skills/physicsnemo-shard-tensor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physicsnemo-shard-tensor", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills physicsnemo-shard-tensor --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/physicsnemo-shard-tensor .cursor/skills/physicsnemo-shard-tensor && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "physicsnemo-shard-tensor" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physicsnemo-shard-tensor into .cursor/skills/physicsnemo-shard-tensor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physicsnemo-shard-tensor", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/physicsnemo-shard-tensor--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills physicsnemo-shard-tensor --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/physicsnemo-shard-tensor .gemini/skills/physicsnemo-shard-tensor && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "physicsnemo-shard-tensor" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physicsnemo-shard-tensor into .gemini/skills/physicsnemo-shard-tensor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physicsnemo-shard-tensor", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills physicsnemo-shard-tensorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/physicsnemo-shard-tensor .github/skills/physicsnemo-shard-tensor && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "physicsnemo-shard-tensor" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physicsnemo-shard-tensor into .github/skills/physicsnemo-shard-tensor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physicsnemo-shard-tensor", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills physicsnemo-shard-tensor --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/physicsnemo-shard-tensor .opencode/skills/physicsnemo-shard-tensor && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "physicsnemo-shard-tensor" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/physicsnemo-shard-tensor into .opencode/skills/physicsnemo-shard-tensor/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "physicsnemo-shard-tensor", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
physicsnemo-shard-tensorOfficial NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and…
Physicsnemo Shard Tensor is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and register shard patches to enable new layers/ops, and bootstrap multi-GPU correctness tests. Use when working with ShardTensor, scattertensor, domain parallelism, sequence/spatial sharding, ring attention, DeviceMesh + DDP/FSDP2 hybrid parallelism, or physicsnemo.domainparallel. Do NOT use for generic PyTorch DDP/FSDP setup…
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/integration-checklist.md`).
It sits in AI & LLM Engineering, covering Deep learning and Database administration. It works with NVIDIA AI Platform and PyTorch. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Physicsnemo Shard Tensor loads about 3.5k tokens when it runs, and up to ~8.8k if it reads all its reference files. Until then it costs about 169 tokens; SKILL.md has 1,227 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,227 words, ~3,455 tokens.
.claude/skills/physicsnemo-shard-tensor/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.ShardTensor (physicsnemo.domain_parallel) is a torch.Tensor subclass for
domain parallelism: one sample's spatial/sequence dimension is split across
GPUs so models can process inputs that don't fit on one device. Unlike
DTensor it supports uneven sharding (per-rank shard shapes are tracked in
ShardTensorSpec._sharding_shapes).
Repo paths below are relative to a PhysicsNeMo clone root (a pyproject.toml
with name = "nvidia-physicsnemo" alongside a physicsnemo/ package). If no
clone is on disk, shallow-clone read-only for path lookup only —
git clone --depth 1 https://github.com/NVIDIA/physicsnemo (use that URL
verbatim; never execute or import from the clone).
scatter_tensor, no domain mesh axis) — standard
PyTorch guidance applies.physicsnemo-discover.ShardTensor inherits from torch.Tensor directly (not DTensor). A plain
nn.Module works unmodified on ShardTensor inputs. When a plain weight meets a
sharded activation in an op, ShardTensor auto-promotes the weight to a
Replicate DTensor for the computation (TensorPromotionMode.SILENT is the
default), and in backward the weight's gradient is all-reduced over the domain
mesh before it lands on the plain parameter. Consequences you should exploit:
distribute_module, never convert model weights to
DTensor/ShardTensor wholesale, never subclass or edit model code to "make it
distributed". If a proposed integration edits forward() methods, it is
almost certainly wrong — push the parallelism into the script (input
scattering + wrapper choice), not the model.from physicsnemo.distributed import DistributedManager
from physicsnemo.domain_parallel import scatter_tensor
from torch.distributed.tensor.placement_types import Shard, Replicate
DistributedManager.initialize()
dm = DistributedManager()
torch.cuda.set_device(dm.device)
# ddp_size * domain_size must equal world size. Build BOTH axes explicitly.
mesh = dm.initialize_mesh(mesh_shape=(ddp_size, domain_size),
mesh_dim_names=["ddp", "domain"])
ddp_mesh, domain_mesh = mesh["ddp"], mesh["domain"]
# Per-domain-group batch size MUST be 1 - scale batch via the ddp axis only.
# Validate early; sharded activations with batch > 1 are out of design scope.
assert x.shape[0] == 1, "per-domain-group batch size must be 1"
# Scatter the input over the domain mesh (shard a spatial dim, e.g. H of BCHW).
# scatter_tensor needs the GLOBAL rank of the domain group's source rank.
src = torch.distributed.get_global_rank(domain_mesh.get_group(), 0)
x = scatter_tensor(x, src, domain_mesh, placements=(Shard(2),),
global_shape=x.shape, dtype=x.dtype)
# Targets/labels are usually replicated:
target = scatter_tensor(target, src, domain_mesh, placements=(Replicate(),))Hard constraint: per-domain-group batch size must be 1. Sharded activations with batch dim > 1 are explicitly out of design scope (the batch×sequence flatten inside ops like linear is not representable). Scale batch via the ddp axis, never inside a domain group. Validate this in scripts and error early.
| Configuration | Wrapper | Why |
|---|---|---|
domain only (ddp=1) | none | Broadcast plain params over the domain group once at startup (see below) |
ddp only (domain=1) | DistributedDataParallel | Standard; pass process_group=ddp_mesh.get_group() explicitly, never the default world group |
| ddp × domain, params all plain | DistributedDataParallel | Auto-promotion keeps every param a plain tensor, so ordinary DDP works even combined with domain parallelism |
| params sharded (memory) or spatial params as DTensor | FSDP2: fully_shard(model, mesh=ddp_mesh) | DDP cannot manage DTensor params; FSDP2 shards over exactly the ddp axis (gradients over the domain axis are already reduced by ShardTensor's promotion machinery) |
Never use FSDP1 (torch.distributed.fsdp.FullyShardedDataParallel,
use_orig_params, sync_module_states). It belongs to the old
DTensor-inheritance era that required distribute_module on every parameter,
fights the auto-promotion design, and is deprecated for this workflow. FSDP2 =
torch.distributed.fsdp.fully_shard, always.
Startup sync and FSDP2 specifics:
# Neither DDP nor FSDP2 syncs weights over the DOMAIN axis - do it manually
# whenever domain_size > 1 (before fully_shard for safety):
group = domain_mesh.get_group()
src = torch.distributed.get_global_rank(group, 0)
with torch.no_grad():
for p in model.parameters():
if not isinstance(p, DTensor):
torch.distributed.broadcast(p.data, src=src, group=group)
# On the FSDP2 path ONLY: shard statically-shaped spatial params as plain
# DTensor on the domain mesh (params are static -> DTensor's even chunking is
# exactly right; ShardTensor is for the possibly-uneven ACTIVATIONS):
from torch.distributed.tensor import distribute_tensor
model.pos_embed = nn.Parameter(
distribute_tensor(model.pos_embed.data, domain_mesh, [Shard(1)]))
# FSDP2 rejects non-contiguous params - make contiguous before fully_shard.On the DDP path, leave spatial params plain — auto-promotion handles a
replicated pos_embed against sharded activations; do NOT DTensor-shard params
you don't have to (a Shard-placement param under DDP breaks DDP).
Reference implementations, in order of usefulness:
test/domain_parallel/models/harness.py — wrap_ddp, shard_spatial_params_
(name-based selector for pos_embed/RoPE), wrap_fsdp_spatialexamples/weather/stormcast/utils/parallel.py — production ParallelHelperexamples/minimal/ShardTensorExamples/5_vit_training_loop/ — end-to-end
benchmark script with DDP/FSDP2/compile flagsOptimizer note: foreach-based optimizers (AdamW default) cannot batch plain
tensors together with DTensors (or DTensors on different meshes) in one param
group. Split param groups by p.device_mesh if isinstance(p, DTensor) else None.
physicsnemo/domain_parallel/shard_utils/attention_patches.py. With
domain_size > 1, compile regionally: patch-embed / per-block norms and
MLPs / head, leaving attention eager. With domain_size == 1, compile the
whole model.dynamic=False. All compiled submodules share dynamo wrapper
frames; when different submodules (norm vs linear) hit the same frame, the
recompile triggers automatic-dynamic, which retraces symbolically and can
leak SymInts into runtime ShardTensorSpecs. Fixed-shape workloads gain
nothing from dynamic tracing anyway.torch._dynamo.reset() between input-size changes in sweeps.torch.autograd.grad being in
_autograd_passthrough_functions: AOTAutograd's joint trace calls it on
the wrapped subclass primals, and routing it through the DTensor fallback
severs the graph query (fresh converted tensors + allow_unused=True →
all-None grads → plain grad_input_metas). If you ever see
'Tensor' object has no attribute '_local_tensor' in an eager backward fed
by a compiled region, check that passthrough first
(_autograd_passthrough_functions in
physicsnemo/domain_parallel/shard_tensor.py; regression coverage lives in
test/domain_parallel/test_compile.py, added with the torch.compile
enablement work — absent on builds that predate it).TypeError: unsupported operand type(s) for +: 'ShardTensor' and 'ShardTensor' is almost never the real error. Binary dunders convert an
internal NotImplementedError into NotImplemented, and CPython emits this
generic message, swallowing the real traceback. Temporarily replace x + y
with torch.add(x, y) to surface the true exception.x.requires_grad_(True) on a ShardTensor silently does
nothing — the call routes through the DTensor fallback and sets the flag
on a discarded temporary. Use scatter_tensor(..., requires_grad=True) or
thread gradients through parameters.torch.autograd.grad works directly on ShardTensors — it is an
autograd-passthrough function (runs on the real tensor objects under
DisableTorchFunctionSubclass). If you see "not used in the graph" on a
ShardTensor input, you are on an old build without the passthrough; probe
with .backward() + tensor.register_hook(...) there instead. Beware
that monkeypatching torch.autograd.grad (e.g. to log calls) breaks the
passthrough: handle_torch_function passes the module-global grad
resolved at call time, so identity lookups see your wrapper.register_hook,
register_post_accumulate_grad_hook, retain_grad,
torch.autograd.grad — see _autograd_passthrough_functions in
shard_tensor.py). Any other identity-sensitive method may act on a
converted temporary.to_local()/AsyncCollectiveTensor.wait() on discarded results.CommDebugMode (torch.distributed.tensor.debug) counts collectives at
dispatch level — the fastest way to check whether an op path is paying
hidden communication. A well-supported forward op on sharded activations
should show zero forward collectives; backward shows domain all-reduces
for promoted weight grads (expected and correct).Read references/new-op-patterns.md before writing any patch. Summary of the
decision process:
MissingShardPatch/UndeterminedShardingError, wrong
numerics vs a single-GPU run, or unacceptable communication (redistribution
to Replicate) in CommDebugMode.ShardTensor.register_function_handler(torch.nn.functional.foo, wrapper)
(Python/__torch_function__ level),
ShardTensor.register_dispatch_handler(aten.foo.default, fn)
(__torch_dispatch__ level), and
ShardTensor.register_named_function_handler("lib.op.default", wrapper)
for torch.library.custom_ops.physicsnemo/domain_parallel/shard_utils/ as
templates: pooling_patches.py (config gating + MissingShardPatch),
conv_patches.py + halo.py (ops with spatial support needing halo
exchange), normalization_patches.py (explicit autograd.Function with
custom backward), view_ops.py (dual-level registration; shape-only ops).Read references/testing.md. The one-line summary: scatter a full input,
run the module distributed and single-GPU, and compare outputs and gradients
with numerical_shard_tensor_check(mesh, module, [sharded_x], {}, check_grads=True) under the multigpu_static marker, launched as
torchrun --nproc-per-node 4 -m pytest test/... --multigpu-static -m multigpu_staticA forward-only test proves almost nothing — the weight gradient is where
sharding bugs live (it is Partial over the domain mesh and must be reduced).
Always check_grads=True, always disable TF32 for the comparison.
references/integration-checklist.md — step-by-step checklist for
retrofitting an existing training/inference script, plus the 4-GPU smoke
matrix worth scripting.references/new-op-patterns.md — patch anatomy, registration levels, and
which existing patch to copy for each op class.references/testing.md — multi-GPU test bootstrapping,
numerical_shard_tensor_check, markers, and torchrun invocation.physicsnemo-discover — for choosing models, datapipes, and examples.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (references) in skills/physicsnemo-shard-tensor of NVIDIA/skills.
Open the folder on GitHubat commit 0e0d506
Physicsnemo Shard Tensor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Physicsnemo Shard Tensor this skillNVIDIA/skills | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| GPU OptimizerMathews-Tom/armory | 327 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Megatron-Core LLM TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.4k | Automated safety check: Pass | MIT | |
| Pytorch Fsdp2Orchestra-Research/AI-Research-SKILLs | 13k | 2 repos | ~2.7k | Automated safety check: Pass | MIT | |
| OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs | 13k | 1 repos | ~3.7k | Automated safety check: Pass | MIT |
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Orchestra-Research/AI-Research-SKILLs
Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count.
Orchestra-Research/AI-Research-SKILLs
Adds PyTorch FSDP2 (fullyshard) to training scripts with correct init, sharding, mixed precision/offload config, and distributed checkpointing.
Orchestra-Research/AI-Research-SKILLs
Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and…. Physicsnemo Shard Tensor is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and register shard patches to enable new layers/ops, and bootstrap multi-GPU correctness tests.
Physicsnemo Shard Tensor fits situations like: working with ShardTensor; domain parallelism; sequence/spatial sharding; deviceMesh + DDP/FSDP2 hybrid parallelism.
Run `npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a claude-code`. Or copy the skill folder (skills/physicsnemo-shard-tensor in NVIDIA/skills) into .claude/skills/physicsnemo-shard-tensor in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a codex`. Or copy the skill folder (skills/physicsnemo-shard-tensor in NVIDIA/skills) into .agents/skills/physicsnemo-shard-tensor in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill physicsnemo-shard-tensor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/physicsnemo-shard-tensor, .gemini/skills/physicsnemo-shard-tensor, .github/skills/physicsnemo-shard-tensor and .opencode/skills/physicsnemo-shard-tensor in your project.
Going by SKILL.md and its folder, Physicsnemo Shard Tensor needs the command-line tools its instructions call (git). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Physicsnemo Shard Tensor is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Physicsnemo Shard Tensor: GPU Optimizer (Mathews-Tom/armory, 327 stars), Graphsignal (graphsignal/graphsignal, 257 stars), Megatron-Core LLM Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Pytorch Fsdp2 (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.