Official agent skill

Nemo Mbridge Resiliency

by NVIDIA in NVIDIA/skills

Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.

OfficialApache-2.0Auto-check passed

Install Nemo Mbridge Resiliency

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-resiliency -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-resiliency --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-resiliency .claude/skills/nemo-mbridge-resiliency && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-resiliency
GitHub stars
3.5k
Token cost
~2.8k tokens
SKILL.md length
648 words
Files
6
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.

  • Works in 7 steps: ft_launcher, not torchrun: Direct… → Async save requires torch_dist:… → IPR + NeMo-Run: In-process restart is… → …
  • SKILL.md covers Enablement, Code Anchors, Pitfalls and Verification
  • Calls uv and pytest

What it does

Nemo Mbridge Resiliency is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

Example prompts

  • “/nemo-mbridge-resiliency”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. ft_launcher, not torchrun: Direct FaultToleranceConfig requires
  2. Async save requires torch_dist: async_save=True only works with
  3. IPR + NeMo-Run: In-process restart is not compatible with NeMo-Run
  4. NVRx vs legacy straggler: Two detectors exist. Use NVRx
  5. stop_if_detected default: NVRx logs but does not stop training by
  6. NCCL watchdog vs hard_timeout: For IPR, NCCL watchdog timeout must
  7. Rerun state machine is alpha: Use check_for_nan_in_loss=True for

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Resiliency loads about 2.8k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 648 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 648 words, ~2,826 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-resiliency/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-resiliency
description
Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.
license
Apache-2.0
when_to_use
Enabling resiliency features, or investigating a commit that caused training hangs, straggler detection failures, or broken restarts; 'fault tolerance'…

Resiliency

Stable docs: @docs/training/resiliency.md, @docs/training/checkpointing.md Card: @skills/nemo-mbridge-resiliency/card.yaml

Enablement

Fault tolerance (Slurm only)
python
from megatron.bridge.recipes.run_plugins import FaultTolerancePlugin
import nemo_run as run

task = run.Script(...)
run_plugins = [
    FaultTolerancePlugin(
        enable_ft_package=True,
        calc_ft_timeouts=True,
        num_in_job_restarts=3,
        num_job_retries_on_failure=2,
        initial_rank_heartbeat_timeout=1800,
        rank_heartbeat_timeout=300,
    )
]
run.run(task, plugins=run_plugins, executor=executor)
Plugin parameterDefaultDescription
num_in_job_restarts3Max restarts within same job
num_job_retries_on_failure2Max new job launches on failure
initial_rank_heartbeat_timeout1800First heartbeat timeout (seconds)
rank_heartbeat_timeout300Subsequent heartbeat timeout (seconds)
Option 2: Direct config + ft_launcher
python
from megatron.bridge.training.config import FaultToleranceConfig

cfg.ft = FaultToleranceConfig(
    enable_ft_package=True,
    calc_ft_timeouts=True,
    simulate_fault=False,
    simulated_fault_type="random",
)

Launch with ft_launcher (not torchrun):

bash
export GROUP_RANK=0  # required for non-Slurm
ft_launcher \
    --rdzv_backend=c10d --rdzv_endpoint=${MASTER_ADDR}:${MASTER_PORT} \
    --nnodes=${NUM_NODES} --nproc-per-node=${NUM_GPUS_PER_NODE} \
    --ft-rank_section_timeouts=setup:600,step:180,checkpointing:420 \
    --ft-rank_out_of_section_timeout=300 \
    your_training_script.py
Config parameterDefaultDescription
enable_ft_packageFalseEnable fault tolerance
calc_ft_timeoutsFalseAuto-compute optimal timeouts
simulate_faultFalseEnable fault simulation for testing
simulated_fault_type"random""rank_hung", "rank_killed", or "random"
simulated_fault_rankNoneSpecific rank to fault (random if None)
simulated_fault_base_delay0Base delay before simulating fault

Section-based timeout monitoring covers setup, training steps, checkpointing, and out-of-section time independently. Timeouts are saved to ft_state.json for subsequent runs when calc_ft_timeouts=True.

NVRx straggler detection
python
from megatron.bridge.training.config import NVRxStragglerDetectionConfig

cfg.nvrx_straggler = NVRxStragglerDetectionConfig(
    enabled=True,
    report_time_interval=300.0,
    calc_relative_gpu_perf=True,
    calc_individual_gpu_perf=True,
    num_gpu_perf_scores_to_print=5,
    gpu_relative_perf_threshold=0.7,
    gpu_individual_perf_threshold=0.7,
    stop_if_detected=False,
    enable_logging=True,
)
ParameterDefaultDescription
enabledFalseEnable straggler detection
report_time_interval300.0Seconds between straggler checks
calc_relative_gpu_perfTrueCompare ranks against each other
calc_individual_gpu_perfTrueTrack per-rank degradation over time
gpu_relative_perf_threshold0.7Threshold for relative performance (0-1)
gpu_individual_perf_threshold0.7Threshold for individual performance (0-1)
stop_if_detectedFalseTerminate training on straggler
num_gpu_perf_scores_to_print5Number of best/worst scores to print
profiling_interval1Profiling interval for detector
Preemption
Plugin (Slurm)
python
from megatron.bridge.recipes.run_plugins import PreemptionPlugin

plugins = [
    PreemptionPlugin(
        preempt_time=60,
        enable_exit_handler=True,
        enable_exit_handler_for_data_loader=False,
    )
]
Plugin parameterDefaultDescription
preempt_time60Seconds before job limit to send signal
enable_exit_handlerTrueEnable signal handler in training
enable_exit_handler_for_data_loaderFalseEnable for dataloader workers
Direct config
python
import signal
cfg.train.exit_signal_handler = True
cfg.train.exit_signal = signal.SIGTERM
cfg.train.exit_signal_handler_for_dataloader = False
Re-run state machine (experimental)
python
from megatron.bridge.training.config import RerunStateMachineConfig

cfg.rerun_state_machine = RerunStateMachineConfig(
    rerun_mode="validate_results",
    check_for_nan_in_loss=True,
    check_for_spiky_loss=False,
    spiky_loss_factor=10.0,
)
ParameterDefaultDescription
rerun_mode"disabled""disabled", "validate_results", "report_determinism_stats"
check_for_nan_in_lossTrueCheck for NaN in loss
check_for_spiky_lossFalseCheck for unexpectedly large loss
spiky_loss_factor10.0Loss flagged if > factor * max observed (increase for large models)

Exit codes: 16 = resume to disambiguate, 17 = failed validation.

In-process restart (experimental)
python
from megatron.bridge.training.config import InProcessRestartConfig

cfg.inprocess_restart = InProcessRestartConfig(
    enabled=True,
    granularity="node",
    soft_timeout=60.0,
    hard_timeout=90.0,
)
ParameterDefaultDescription
enabledFalseEnable in-process restart
active_world_sizeNoneRanks executing workload (rest are warm reserves)
granularity"node""node" or "rank" restart granularity
max_iterationsNoneMax restart attempts (None = unlimited)
soft_timeout60.0Detect GIL-released hangs (seconds)
hard_timeout90.0Force-terminate hung ranks (seconds)
heartbeat_interval30.0Heartbeat interval (seconds)
heartbeat_timeout60.0Missing heartbeat timeout (seconds)
barrier_timeout120.0Distributed barrier timeout (seconds)
completion_timeout120.0Completion barrier timeout (seconds)
empty_cuda_cacheTrueClear CUDA cache during restart
max_rank_faultsNoneMax rank faults before terminating
monitor_process_logdirNoneDirectory for monitor logs

Required environment variables:

bash
export TORCH_CPP_LOG_LEVEL=error
export TORCH_NCCL_RETHROW_CUDA_ERRORS=0
export NCCL_NVLS_ENABLE=0

The PyTorch NCCL watchdog timeout must exceed hard_timeout. NeMo-Run's Slurm Executor is not supported; launch directly with srun --kill-on-bad-exit=0.

Async checkpoint save
python
cfg.checkpoint.async_save = True
cfg.checkpoint.ckpt_format = "torch_dist"
Local checkpointing (NVRx)
python
cfg.checkpoint.non_persistent_local_ckpt_dir = "/local/scratch/ckpt"
cfg.checkpoint.non_persistent_local_ckpt_algo = "fully_parallel"

Code Anchors

Show full SKILL.md (262 more words)Show less
Fault tolerance
  • Config: src/megatron/bridge/training/config.py — FaultToleranceConfig
  • Runtime: src/megatron/bridge/training/fault_tolerance.py
  • Plugin: src/megatron/bridge/recipes/run_plugins.py — FaultTolerancePlugin
  • Perf plugin: scripts/performance/nemo-mbridge-resiliency_plugins.py
  • Tests: tests/unit_tests/training/test_fault_tolerance.py
  • Example: examples/training_features/nemo-mbridge-resiliency/fault_tolerance/
Straggler detection
  • Config: src/megatron/bridge/training/config.py — NVRxStragglerDetectionConfig
  • Runtime: src/megatron/bridge/training/nvrx_straggler.py
  • Train loop: src/megatron/bridge/training/train.py — check_nvrx_straggler_detection
  • Tests: tests/unit_tests/training/test_nvrx_straggler.py, tests/functional_tests/training/test_nvrx_straggler.py
  • Example: examples/training_features/nemo-mbridge-resiliency/straggler_detection/
In-process restart
  • Config: src/megatron/bridge/training/config.py — InProcessRestartConfig
  • Runtime: src/megatron/bridge/training/inprocess_restart.py
  • Entry point: src/megatron/bridge/training/pretrain.py — maybe_wrap_for_inprocess_restart
  • Tests: tests/unit_tests/training/test_inprocess_restart.py, tests/functional_tests/training/test_inprocess_restart.py
Preemption
  • Plugin: src/megatron/bridge/recipes/run_plugins.py — PreemptionPlugin
  • Signal handler: src/megatron/bridge/training/utils/sig_utils.py
  • Tests: tests/unit_tests/recipes/test_run_plugins.py
Re-run state machine
  • Config: src/megatron/bridge/training/config.py — RerunStateMachineConfig
  • Init: src/megatron/bridge/training/initialize.py — init_rerun_state
Checkpointing
  • Async save: src/megatron/bridge/training/checkpointing.py — schedule_async_save
  • Local ckpt: src/megatron/bridge/training/checkpointing.py — LocalCheckpointManager
  • Tests: tests/functional_tests/training/test_local_checkpointing.py

Pitfalls

  1. ft_launcher, not torchrun: Direct FaultToleranceConfig requires ft_launcher. Using torchrun silently disables FT. For non-Slurm, set GROUP_RANK=0.

  2. Async save requires torch_dist: async_save=True only works with ckpt_format="torch_dist". Other formats silently fail or error.

  3. IPR + NeMo-Run: In-process restart is not compatible with NeMo-Run or Slurm preemption plugins. Requires specific PyTorch/NCCL versions and env vars.

  4. NVRx vs legacy straggler: Two detectors exist. Use NVRx (nvrx_straggler); do not enable both.

  5. stop_if_detected default: NVRx logs but does not stop training by default. Set stop_if_detected=True for automatic termination.

  6. NCCL watchdog vs hard_timeout: For IPR, NCCL watchdog timeout must exceed hard_timeout or PyTorch kills the process before recovery.

  7. Rerun state machine is alpha: Use check_for_nan_in_loss=True for NaN detection, but don't rely on full rerun workflows yet.

Verification

Fault tolerance
bash
./examples/training_features/nemo-mbridge-resiliency/fault_tolerance/run_fault_tolerance.sh
./examples/training_features/nemo-mbridge-resiliency/fault_tolerance/run_fault_tolerance.sh --simulate-fault

Look for [FaultTolerance] / [RankMonitorServer] log lines with section timeouts. Simulated fault should trigger restart from checkpoint.

Straggler detection
bash
uv run python -m torch.distributed.run --nproc_per_node=2 \
    examples/training_features/nemo-mbridge-resiliency/straggler_detection/straggler_detection_example.py

Look for GPU relative performance and GPU individual performance reports with per-rank scores.

Async checkpoint

Look for Scheduling async checkpoint save in logs. Training iterations should continue while checkpoint files are being written.

In-process restart
bash
pytest tests/functional_tests/training/test_inprocess_restart.py -v

Requires compatible PyTorch/NCCL versions.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-resiliency of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemo Mbridge Resiliency next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Resiliency compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Resiliency this skillNVIDIA/skills3.5k—~2.8kAutomated safety check: PassApache-2.0
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Skill InspectorNVIDIA/SkillSpector20k—~1.8kAutomated safety check: PassApache-2.0
Megatron-LM Container and Dependency SetupNVIDIA/Megatron-LM18k—~2.6kAutomated safety check: PassApache-2.0
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
Megatron-LM Base Image BumpNVIDIA/Megatron-LM18k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Skill Inspector

    NVIDIA/SkillSpector

    Official

    Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.

    20k GitHub stars~1.8k tokensUpdated today
    SecurityAuto-check passed
  • Official

    Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.

    18k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Remove bracketed NemoClaw tags from GitHub issue and PR titles.

    23k GitHub stars~693 tokensUpdated today
    Marketing & SEOAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Nemo Mbridge Resiliency

What does Nemo Mbridge Resiliency do?

Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine. Nemo Mbridge Resiliency is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Resiliency features in Megatron Bridge including fault tolerance, straggler detection, in-process restart, preemption, and re-run state machine.

How do I install Nemo Mbridge Resiliency in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-resiliency -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-resiliency in NVIDIA/skills) into .claude/skills/nemo-mbridge-resiliency in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Resiliency in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-resiliency -a codex`. Or copy the skill folder (skills/nemo-mbridge-resiliency in NVIDIA/skills) into .agents/skills/nemo-mbridge-resiliency in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Resiliency in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-resiliency -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-resiliency, .gemini/skills/nemo-mbridge-resiliency, .github/skills/nemo-mbridge-resiliency and .opencode/skills/nemo-mbridge-resiliency in your project.

What does Nemo Mbridge Resiliency need to run?

Going by SKILL.md and its folder, Nemo Mbridge Resiliency needs the command-line tools its instructions call (uv and pytest). Our summary lists: Python 3.

Does Nemo Mbridge Resiliency access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Resiliency safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Resiliency use?

Nemo Mbridge Resiliency is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Resiliency use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Resiliency?

Skills that share tags, products or a category with Nemo Mbridge Resiliency: LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Skill Inspector (NVIDIA/SkillSpector, 20k stars), Megatron-LM Container and Dependency Setup (NVIDIA/Megatron-LM, 18k stars) and Embeddings via 9Router (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Resiliency?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.