DGX Spark Training Gotchas
wshobson/agents
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
$ npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-automodel-distributed-training --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-automodel-distributed-training .claude/skills/nemo-automodel-distributed-training && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-automodel-distributed-training" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-distributed-training into .claude/skills/nemo-automodel-distributed-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-automodel-distributed-training", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-distributed-trainingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-automodel-distributed-training --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-automodel-distributed-training .agents/skills/nemo-automodel-distributed-training && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-automodel-distributed-training" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-distributed-training into .agents/skills/nemo-automodel-distributed-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-automodel-distributed-training", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-automodel-distributed-training --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-automodel-distributed-training .cursor/skills/nemo-automodel-distributed-training && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-automodel-distributed-training" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-distributed-training into .cursor/skills/nemo-automodel-distributed-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-automodel-distributed-training", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-automodel-distributed-training--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-automodel-distributed-training --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-automodel-distributed-training .gemini/skills/nemo-automodel-distributed-training && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-automodel-distributed-training" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-distributed-training into .gemini/skills/nemo-automodel-distributed-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-automodel-distributed-training", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-automodel-distributed-trainingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-automodel-distributed-training .github/skills/nemo-automodel-distributed-training && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-automodel-distributed-training" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-distributed-training into .github/skills/nemo-automodel-distributed-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-automodel-distributed-training", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-automodel-distributed-training --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-automodel-distributed-training .opencode/skills/nemo-automodel-distributed-training && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-automodel-distributed-training" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-automodel-distributed-training into .opencode/skills/nemo-automodel-distributed-training/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-automodel-distributed-training", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-automodel-distributed-trainingGuide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
Nemo Automodel Distributed Training is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `BENCHMARK.md`, `evals/evals.json` and `skill-card.md`).
It sits in AI & LLM Engineering, covering Deep learning. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and python).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Automodel Distributed Training loads about 4.8k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 1,487 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,487 words, ~4,816 tokens.
.claude/skills/nemo-automodel-distributed-training/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.NeMo AutoModel uses PyTorch-native distributed training.
All parallelism is orchestrated through a single MeshContext object that
holds device meshes, strategy configs, and axis names.
<!-- NVSkills catalog signing requested after PR #2937 (2026-07-31). -->
For conceptual distributed-training questions, answer directly from the quick patterns in this skill without inspecting the repository. Start with the strategy choice, then list only the YAML fields and constraints relevant to the question.
Use direct action verbs in the final answer: recommend the strategy, show the minimal YAML, state the sizing constraint, and name the unsupported strategies. Do not discuss model onboarding, recipes, Slurm, SkyPilot, or checkpointing unless the user asks.
Recommend strategy: fsdp2. Mention tp_size, pp_size, cp_size,
ep_size, and the pipeline sub-config. State that dp_size is inferred from
world_size / (tp_size * pp_size * cp_size).
distributed:
strategy: fsdp2
tp_size: 8
pp_size: 4
cp_size: 1
ep_size: 1
pipeline:
pp_schedule: interleaved1f1b
pp_microbatch_size: 1Recommend strategy: fsdp2 with ep_size > 1. Say this creates a separate
moe_mesh; include the moe sub-config when relevant; state that ep_size
must divide dp_size * cp_size. Do not recommend megatron_fsdp or ddp.
distributed:
strategy: fsdp2
ep_size: 8
moe:
reshard_after_forward: falseSay no for pipeline parallelism, expert parallelism, and sequence_parallel.
Recommend fsdp2 for PP, EP, or sequence_parallel; mention that DDP is only
simple data parallelism.
Three strategies are available, selected via the distributed.strategy YAML key:
| Strategy | YAML value | Best for |
|---|---|---|
| FSDP2 | fsdp2 | General use, recommended default. Supports TP, PP, CP, EP, HSDP. |
| MegatronFSDP | megatron_fsdp | NVIDIA Megatron-style FSDP. No PP, no EP, no sequence_parallel. |
| DDP | ddp | Simple data parallelism only. No TP, PP, CP, or EP. |
Decision tree:
fsdp2 (default). Use ddp only if you need the simplest possible setup.fsdp2 with appropriate TP/PP sizing.fsdp2 with ep_size > 1 (creates a separate moe_mesh).fsdp2 with PP + TP.cp_size > 1).When answering strategy-selection questions, state the chosen distributed.strategy
first, then enumerate the YAML fields the user must set.
Quick TP + PP answer:
strategy: fsdp2; do not use megatron_fsdp when pipeline parallelism is required.tp_size for tensor parallelism and pp_size for pipeline parallelism.pipeline: sub-config with pp_schedule and pp_microbatch_size.dp_size unset or none; it is inferred as world_size / (tp_size * pp_size * cp_size).Quick MoE expert-parallel answer:
strategy: fsdp2 and ep_size > 1.moe: sub-config only when ep_size > 1; it maps to MoEParallelizerConfig.moe_mesh for expert parallelism in addition to the main device_mesh.megatron_fsdp or ddp for expert parallelism; megatron_fsdp has no EP support.ep_size must divide dp_size * cp_size and that megatron_fsdp does not support EP, PP, or sequence_parallel.The distributed section in the recipe YAML maps directly to
parse_distributed_section() in recipes/_dist_utils.py:
distributed:
strategy: fsdp2 # fsdp2 | megatron_fsdp | ddp
dp_size: none # auto-calculated from world_size / (tp * pp * cp)
dp_replicate_size: none # FSDP2-only, for HSDP
tp_size: 1
pp_size: 1
cp_size: 1
ep_size: 1
# Strategy-specific flags (forwarded to the strategy dataclass):
sequence_parallel: false
activation_checkpointing: false
defer_fsdp_grad_sync: true # FSDP2 only
# Sub-configs (optional):
pipeline:
pp_schedule: 1f1b
pp_microbatch_size: 1
# ... see PipelineConfig fields
moe:
reshard_after_forward: false
# ... see MoEParallelizerConfig fieldsThe dp_size is always inferred:
dp_size = world_size / (tp_size * pp_size * cp_size)initialize_distributed() [components/distributed/init_utils.py]
-> initializes torch.distributed process group and returns DistInfo
YAML distributed section + DistInfo.world_size
-> parse_distributed_section() [recipes/_dist_utils.py]
-> create_distributed_setup_from_config() [recipes/_dist_utils.py]
-> DistributedSetup.build() [components/distributed/config.py]
-> instantiate_infrastructure() [_transformers/infrastructure.py]
-> _instantiate_distributed() -> FSDP2Manager / MegatronFSDPManager / DDPManager
-> _instantiate_pipeline() -> AutoPipeline (if pp_size > 1)
-> parallelize_fn -> MoE parallelizer (if ep_size > 1) or PP wrapper
-> apply_model_infrastructure() [_transformers/infrastructure.py]
-> _shard_pp() or _shard_ep_fsdp() (applies sharding to the model)distributed:
strategy: fsdp2
tp_size: 1
cp_size: 1This auto-calculates dp_size = world_size and applies fully_shard() per
transformer block via DTensor-based sharding.
Keep TP within a single NVLink domain (typically one node):
distributed:
strategy: fsdp2
tp_size: 4 # 2, 4, or 8 -- must divide GPUs per node
sequence_parallel: trueThe TP plan is auto-selected based on the model type. Pass a custom plan via the Python API if needed:
config = FSDP2Config(sequence_parallel=True, tp_plan=my_custom_plan)distributed:
strategy: fsdp2
pp_size: 2
pipeline:
pp_schedule: interleaved1f1b # 1f1b, gpipe, interleaved_1f1b, etc.
pp_microbatch_size: 4
scale_grads_in_schedule: falseThe model must have a _pp_plan attribute (set on the HF model class) for
AutoPipeline to know how to split layers across stages. Models without
_pp_plan are not compatible with PP.
Intra-node full sharding + inter-node replication via a 2D DeviceMesh:
distributed:
strategy: fsdp2
dp_replicate_size: 2 # must divide dp_sizeConstraint: dp_replicate_size < dp_size (pure replication with no sharding
is not supported by FSDP2).
Trades compute for memory by recomputing activations during backward:
distributed:
activation_checkpointing: trueThis is a model-build/training behavior flag, not mesh topology. Dense strategies read it from the strategy config; EP/MoE paths pass the recipe-level flag directly into model infrastructure.
FSDP2 defers gradient sync to the final micro-batch by default for communication overlap:
distributed:
defer_fsdp_grad_sync: true # defaultFSDP2Config defaults to bfloat16 for all three precision knobs via
MixedPrecisionPolicy(param_dtype=bf16, reduce_dtype=bf16, output_dtype=bf16, cast_forward_inputs=True). Override via the Python API:
from torch.distributed.fsdp import MixedPrecisionPolicy
config = FSDP2Config(
mp_policy=MixedPrecisionPolicy(param_dtype=torch.float16, reduce_dtype=torch.float32),
)_pp_plan (a dict mapping module FQNs to stages).pp_size > 1 in the distributed section.pipeline sub-config with schedule and microbatch size.Defined in PipelineConfig.pp_schedule:
1f1b (one-forward-one-backward, default)gpipeinterleaved_1f1b / interleaved1f1blooped_bfsdfsv_schedulezero_bubbledistributed:
strategy: fsdp2
pp_size: 2
pipeline:
pp_schedule: interleaved1f1b
pp_microbatch_size: 4
scale_grads_in_schedule: false
checkpoint:
model_save_format: safetensors
save_consolidated: finalAutoPipeline.build() calls pipeline_model() which splits the model into
stages using the model's _pp_plan, creates PipelineStage objects, and
builds the schedule. During training, schedule.step() drives forward and
backward through the pipeline.
Use CP for long sequences (8K+). CP shards Q/K/V on the sequence dimension as DTensors.
distributed:
strategy: fsdp2
cp_size: 2 # or 4, 8is_causal=True is set via
forward pre-hooks registered by attach_context_parallel_hooks().apply_model_infrastructure() calls
attach_context_parallel_hooks() on each model part (for non-TE models).make_cp_batch_and_ctx() creates a CP context
manager that shards the batch along the sequence dimension and sets up
context_parallel() from torch.distributed.tensor.experimental.make_cp_batch_for_te() uses THD format and
TE's thd_get_partitioned_indices for sharding.CP works with packed sequences. The packed_sequence_size must be divisible
by cp_size. When using TE, chunks are sharded per-chunk via
_shard_thd_chunk_for_te().
Packing multiple sequences into a single training sample for efficiency.
packed_sequence:
packed_sequence_size: 4096 # 0 = disabled
step_scheduler:
local_batch_size: 1 # must be 1 for packed sequencesWhen packed_sequence_size > 0, the dataset collator packs sequences up to
that length. local_batch_size must be 1 because each "sample" is already a
packed batch.
Set ep_size > 1 to distribute experts across GPUs. This creates a separate
moe_mesh alongside the main device_mesh:
distributed:
strategy: fsdp2
ep_size: 8
activation_checkpointing: trueThe moe_mesh shape is (pp_size, ep_shard_size, ep_size) with dimension
names ("pp", "ep_shard", "ep").
Constraint: dp_cp_size (= dp_size * cp_size) must be divisible by
ep_size.
distributed:
strategy: fsdp2
ep_size: 8
activation_checkpointing: true
moe:
reshard_after_forward: false
ignore_router_for_ac: false
wrap_outer_model: trueThe moe sub-section maps to MoEParallelizerConfig and is only
instantiated when ep_size > 1.
distributed:
strategy: fsdp2
tp_size: 1
cp_size: 1
pp_size: 1
ep_size: 8
sequence_parallel: false
activation_checkpointing: trueDespite its name, megatron_fsdp does not support expert parallelism
(ep_size > 1), pipeline parallelism (pp_size > 1), or
sequence_parallel. Use fsdp2 for these features.
| Model size | TP | PP | CP | Strategy |
|---|---|---|---|---|
| < 3B | 1 | 1 | 1 | FSDP2 (DP only) |
| 3-13B | 2-4 | 1 | 1 | FSDP2 + TP |
| 13-70B | 4-8 | 2-4 | 1 | FSDP2 + TP + PP |
| 70B+ | 8 | 4-8 | 1 | FSDP2 + TP + PP |
| Any + long seq (8K+) | as above | as above | 2-8 | add CP |
MoE models need less TP than dense models of similar total parameter count because only a fraction of parameters are active per token. EP is the primary scaling dimension:
| Model | TP | PP | EP | Notes |
|---|---|---|---|---|
| Small MoE (<10B total) | 1 | 1 | 8 | EP only |
| Medium MoE (10-30B total) | 1-2 | 1 | 8 | small TP for shared layers |
| Large MoE (100B+ total) | 1-2 | 4+ | 8-64 | PP for depth, EP for experts |
When not using YAML recipes, configure distributed training via Python:
from nemo_automodel.components.distributed import (
DistributedSetup,
FSDP2Config,
ParallelismSizes,
initialize_distributed,
)
dist_env = initialize_distributed("nccl")
distributed_setup = DistributedSetup.build(
strategy=FSDP2Config(sequence_parallel=True),
parallelism_sizes=ParallelismSizes(tp_size=2),
activation_checkpointing=True,
world_size=dist_env.world_size,
)Or pass directly to from_pretrained:
from nemo_automodel import NeMoAutoModelForCausalLM
model = NeMoAutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.2-1B",
distributed_setup=distributed_setup,
)Strategy config dataclasses:
components/distributed/config.py
FSDP2Config -- sequence_parallel, tp_plan, mp_policy, offload_policy,
activation_checkpointing, defer_fsdp_grad_sync
MegatronFSDPConfig -- zero_dp_strategy, overlap_grad_reduce, overlap_param_gather, etc.
DDPConfig -- activation_checkpointing onlyMeshContext (single source of truth for parallelism):
components/distributed/mesh.py
MeshContext -- device_mesh, moe_mesh
Properties: tp_size, pp_size, cp_size, ep_size, dp_size, dp_replicate_size
MeshAxisName -- PP, DP, DP_REPLICATE, DP_SHARD, DP_SHARD_CP, DP_CP, CP, TP, EP, EP_SHARDMesh context and raw mesh creation:
components/distributed/config.py
DistributedSetup.build() -- builds MeshContext from strategy + parallelism
components/distributed/mesh_utils.py
_create_device_meshes() -- routes to FSDP2/MegatronFSDP/DDP raw mesh creation
_create_fsdp2_device_mesh() -- shape (pp, dp_replicate, dp_shard, cp, tp) + flattened submeshes
_create_megatron_fsdp_device_mesh() -- shape (dp, cp, tp)Distributed managers:
components/distributed/fsdp2.py -- FSDP2Manager.parallelize()
components/distributed/megatron_fsdp.py -- MegatronFSDPManager.parallelize()
components/distributed/ddp.py -- DDPManagerPipeline parallelism:
components/distributed/pipelining/config.py -- PipelineConfig dataclass
components/distributed/pipelining/autopipeline.py -- AutoPipeline orchestrator
components/distributed/pipelining/functional.py -- pipeline_model(), schedule creation
components/distributed/pipelining/hf_utils.py -- HF model validation for PPContext parallelism:
components/distributed/context_parallel/utils.py
make_cp_batch_and_ctx() -- creates CP context manager + shards batch
create_context_parallel_ctx() -- wraps torch.distributed.tensor.experimental.context_parallel
attach_context_parallel_hooks() -- strips attention_mask, sets is_causal=True
make_cp_batch_for_te() -- TE-specific CP batch sharding (THD format)Infrastructure orchestration:
_transformers/infrastructure.py
instantiate_infrastructure() -- config objects -> runtime objects
apply_model_infrastructure() -- applies sharding, PEFT, checkpoints to model
_shard_pp() -- pipeline parallel path
_shard_ep_fsdp() -- EP + FSDP path (non-PP)YAML parsing:
recipes/_dist_utils.py
parse_distributed_section() -- YAML dict -> typed configs + sizes
create_distributed_setup_from_config() -- recipe adapter: parse + create DistributedSetup; does not init process groupMoE config:
components/distributed/config.py
MoEParallelizerConfig -- reshard_after_forward, ignore_router_for_ac, wrap_outer_model, etc.
components/moe/config.py
MoEConfig -- n_routed_experts, n_activated_experts, score_func, etc.TP across nodes destroys throughput. Always keep TP within a single NVLink domain. Use PP or DP for cross-node scaling.
PP requires _pp_plan on the model class. Not all HF models have this.
Check validate_hf_model_for_pipeline_support() before enabling PP.
PP bubbles reduce GPU utilization. Use interleaved schedules
(interleaved_1f1b) and smaller microbatches to reduce bubble time.
FSDP2 requires DTensor-aware state dict saving. Use safetensors with
save_consolidated: final for final HF export, or save_consolidated: false
plus the generated model/consolidate.sh helper for offline export.
CP requires compatible attention. SDPA (Flash Attention or Efficient
Attention) or TE attention only. SDPBackend.MATH is not compatible with
DTensor.
MoE EP size must evenly divide dp_size * cp_size. The device mesh
creation asserts dp_cp_size % ep_size == 0.
MegatronFSDP is more limited than FSDP2. It does not support PP
(pp_size > 1), EP (ep_size > 1), or sequence_parallel. The
MeshContext validation raises on these combinations.
DDP supports nothing beyond data parallelism. No TP, PP, CP, EP, or HSDP. Validation raises on any of these.
Activation checkpointing increases compute. It saves memory by recomputing activations during backward, but adds ~30% compute overhead.
Mixed precision policy must match model expectations. The default
bfloat16 policy works for most models. FP16 models may need a custom
MixedPrecisionPolicy.
packed_sequence_size must be divisible by cp_size when using CP
with packed sequences.
dp_replicate_size is FSDP2-only. Passing it with megatron_fsdp
or ddp raises a ValueError.
Run the smallest recipe that exercises the requested strategy. Success means exit code 0, finite loss, no NCCL timeout, and log output matching the expected TP/PP/CP/EP sizes.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in skills/nemo-automodel-distributed-training of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Nemo Automodel Distributed Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Automodel Distributed Training this skillNVIDIA/skills | 3.5k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | |
| DGX Spark Training Gotchaswshobson/agents | 40k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Megatron-LM on SLURMNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs | 13k | 1 repos | ~3.7k | Automated safety check: Pass | MIT | |
| Cosmos Policy EvaluationOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.7k | Automated safety check: Pass | MIT | |
| GPU OptimizerMathews-Tom/armory | 328 | — | ~3.5k | Automated safety check: Notes | MIT |
wshobson/agents
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
Orchestra-Research/AI-Research-SKILLs
Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.
Orchestra-Research/AI-Research-SKILLs
Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings. Nemo Automodel Distributed Training is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Guide for selecting and configuring distributed training strategies in NeMo AutoModel, including FSDP2, Megatron FSDP, DDP, and parallelism settings.
Nemo Automodel Distributed Training fits situations like: tasks that involve Deep learning.
Run `npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a claude-code`. Or copy the skill folder (skills/nemo-automodel-distributed-training in NVIDIA/skills) into .claude/skills/nemo-automodel-distributed-training in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a codex`. Or copy the skill folder (skills/nemo-automodel-distributed-training in NVIDIA/skills) into .agents/skills/nemo-automodel-distributed-training in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-automodel-distributed-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-automodel-distributed-training, .gemini/skills/nemo-automodel-distributed-training, .github/skills/nemo-automodel-distributed-training and .opencode/skills/nemo-automodel-distributed-training in your project.
SKILL.md names no scripts, command-line tools or credentials: Nemo Automodel Distributed Training is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Automodel Distributed Training is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Automodel Distributed Training: DGX Spark Training Gotchas (wshobson/agents, 40k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), OpenVLA-OFT Fine-Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Cosmos Policy Evaluation (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.