Official agent skill

Nemo Mbridge Mlm Bridge Training

by NVIDIA in NVIDIA/skills

Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Mlm Bridge Training

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-mlm-bridge-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-mlm-bridge-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-mlm-bridge-training .claude/skills/nemo-mbridge-mlm-bridge-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-mlm-bridge-training
GitHub stars
3.5k
Token cost
~1.6k tokens
SKILL.md length
391 words
Files
6
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data.

  • Works in 5 steps: Bridge recipe:… → Bridge entry point:… → MLM entry point:… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers First Answer Checklist, Correlation Testing, Multi-GPU Examples and Available Recipes, plus 2 more sections
  • Calls uv and git

What it does

Nemo Mbridge Mlm Bridge Training is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. Covers correlation testing, available recipes, and multi-GPU examples.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/nemo-mbridge-mlm-bridge-training”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Bridge recipe: vanilla_gpt_pretrain_config.
  2. Bridge entry point: scripts/training/run_recipe.py.
  3. MLM entry point: 3rdparty/Megatron-LM/pretrain_gpt.py.
  4. Launch wrapper for both: uv run python -m torch.distributed.run.
  5. Fresh-run cleanup: rm -rf nemo_experiments before the Bridge run.

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Mlm Bridge Training loads about 1.6k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 391 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 391 words, ~1,571 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-mlm-bridge-training/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-mlm-bridge-training
description
Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. Covers correlation testing, available recipes, and multi-GPU examples.
license
Apache-2.0
when_to_use
Running training, comparing MLM vs Bridge loss curves, translating MLM CLI args to Bridge config, or investigating why loss curves diverged after a commit…

MLM vs Bridge Training

For how they differ, the arg mapping tables, gotchas, and translation script, see:

  • @docs/megatron-lm-to-megatron-bridge.md

First Answer Checklist

For MLM-vs-Bridge correlation questions, always name these items up front:

  1. Bridge recipe: vanilla_gpt_pretrain_config.
  2. Bridge entry point: scripts/training/run_recipe.py.
  3. MLM entry point: 3rdparty/Megatron-LM/pretrain_gpt.py.
  4. Launch wrapper for both: uv run python -m torch.distributed.run.
  5. Fresh-run cleanup: rm -rf nemo_experiments before the Bridge run.

Also state that MLM needs PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH, matched Bridge and MLM losses should agree within BF16 rounding, and files under 3rdparty/Megatron-LM/ should not be modified from this repo.

Correlation Testing

Use vanilla_gpt_pretrain_config for loss-correlation testing. This recipe uses bare GPTModelProvider defaults (LayerNorm, GeLU, learned_absolute position embeddings, vocab_size inherited from tokenizer) — matching MLM pretrain_gpt.py defaults with no args.

MLM Correlation Run (2L/256H, 1 GPU)
bash
PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH \
uv run python -m torch.distributed.run --nproc_per_node=1 \
  3rdparty/Megatron-LM/pretrain_gpt.py \
  --num-layers 2 --hidden-size 256 --num-attention-heads 4 \
  --ffn-hidden-size 1024 --seq-length 512 --max-position-embeddings 512 \
  --micro-batch-size 4 --global-batch-size 32 \
  --train-iters 10 --eval-iters 2 --eval-interval 10 \
  --mock-data --bf16 --use-mcore-models \
  --tokenizer-type NullTokenizer --vocab-size 32000 \
  --lr 3e-4 --min-lr 3e-5 --seed 1234 --log-interval 1
Bridge Correlation Run (same config, 1 GPU)
bash
rm -rf nemo_experiments && \
uv run python -m torch.distributed.run --nproc_per_node=1 \
  scripts/training/run_recipe.py \
  --recipe vanilla_gpt_pretrain_config \
  model.num_layers=2 model.hidden_size=256 \
  model.num_attention_heads=4 model.ffn_hidden_size=1024 \
  model.seq_length=512 dataset.seq_length=512 \
  train.train_iters=10 train.global_batch_size=32 train.micro_batch_size=4 \
  validation.eval_interval=10 validation.eval_iters=2 \
  optimizer.lr=3e-4 optimizer.min_lr=3e-5 \
  scheduler.lr_warmup_iters=1 scheduler.lr_decay_iters=10 \
  rng.seed=1234 logger.log_interval=1
Verification

With matched parameters the LM losses should be nearly identical at each iteration. Compare lm loss values from both logs — they should agree to within BF16 rounding.

Multi-GPU Examples

MLM 2-GPU with TP=2
bash
PYTHONPATH=3rdparty/Megatron-LM:$PYTHONPATH \
uv run python -m torch.distributed.run --nproc_per_node=2 \
  3rdparty/Megatron-LM/pretrain_gpt.py \
  --tensor-model-parallel-size 2 --sequence-parallel \
  --num-layers 4 --hidden-size 256 --num-attention-heads 4 \
  --seq-length 1024 --max-position-embeddings 1024 \
  --micro-batch-size 2 --global-batch-size 16 \
  --train-iters 10 --eval-iters 2 --eval-interval 10 \
  --mock-data --bf16 --use-mcore-models \
  --tokenizer-type NullTokenizer --vocab-size 1024 \
  --lr 1e-4 --log-interval 1
Bridge 2-GPU with TP=2
bash
rm -rf nemo_experiments && \
uv run python -m torch.distributed.run --nproc_per_node=2 \
  scripts/training/run_recipe.py \
  --recipe vanilla_gpt_pretrain_config \
  model.tensor_model_parallel_size=2 model.sequence_parallel=true \
  model.num_layers=4 model.hidden_size=256 \
  model.num_attention_heads=4 model.ffn_hidden_size=1024 \
  model.seq_length=1024 dataset.seq_length=1024 \
  train.train_iters=10 train.global_batch_size=16 train.micro_batch_size=2 \
  validation.eval_interval=10 validation.eval_iters=2 \
  scheduler.lr_warmup_iters=2 scheduler.lr_decay_iters=10 \
  logger.log_interval=1

Available Recipes

Common recipes (use with --recipe):

  • vanilla_gpt_pretrain_config — Minimal GPT (bare GPTModelProvider defaults, ideal for correlation testing and custom configs)
  • llama32_1b_pretrain_config — Llama 3.2 1B (16L, 2048H, GBS=512, seq=8192)
  • llama3_8b_pretrain_config — Llama 3 8B
  • qwen3_8b_pretrain_config — Qwen3 8B
  • deepseek_v2_lite_pretrain_config — DeepSeek-V2-Lite 16B MoE

SFT/PEFT variants use _sft_config / _peft_config suffix.

Megatron-Core Submodule

For what the submodule is and why two versions exist, see @docs/megatron-lm-to-megatron-bridge.md.

Check current version
bash
./scripts/switch_mcore.sh status
Show full SKILL.md (157 more words)Show less
Switch to dev for testing newer MCore features
bash
./scripts/switch_mcore.sh dev

# uv sync (without --locked) since lockfile is for main
uv sync
Switch back to main
bash
./scripts/switch_mcore.sh main
After pulling latest main

When you pull the latest Bridge main branch, the submodule pointer may have been updated. Re-sync the submodule:

bash
git submodule update --init 3rdparty/Megatron-LM

Pitfalls

  1. Always rm -rf nemo_experiments before a fresh correlation run. Bridge auto-resumes from stale checkpoints silently.

  2. uv run required: Always use uv run python -m torch.distributed.run (not bare torchrun or python).

  3. MLM PYTHONPATH: Must include 3rdparty/Megatron-LM so gpt_builders.py is importable.

  4. Scheduler overrides: When overriding train.train_iters to a small value, also set scheduler.lr_warmup_iters and scheduler.lr_decay_iters or you get an assertion error.

  5. Use dataset.seq_length in CLI overrides for both pretraining and fine-tuning datasets.

  6. MoE OOM: Large MoE models require full activation recomputation and typically multi-node EP. TP does NOT reduce per-GPU expert memory.

  7. uv sync --locked fails after switching to dev: The lockfile is generated against the main MCore commit. Use uv sync (without --locked) when on dev.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-mlm-bridge-training of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemo Mbridge Mlm Bridge Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Mlm Bridge Training compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Mlm Bridge Training this skillNVIDIA/skills3.5k—~1.6kAutomated safety check: PassApache-2.0
Fla Triton To Gluonfla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
DGX Spark Memory and Thermal Opswshobson/agents40k—~2kAutomated safety check: PassMIT
Yolo Detection 2026SharpAI/DeepCamera3.1k—~1.5kAutomated safety check: PassMIT
Sglang Diffusion Modelopt Quantsgl-project/sglang37k2 repos~5kAutomated safety check: PassApache-2.0

Similar skills

  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub stars~2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Yolo Detection 2026

    SharpAI/DeepCamera

    YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.

    3.1k GitHub stars~1.5k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

    37k GitHub starsUsed in 2 repos~5k tokens
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub stars~2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemo Mbridge Mlm Bridge Training

What does Nemo Mbridge Mlm Bridge Training do?

Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data. Nemo Mbridge Mlm Bridge Training is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run Megatron-LM (MLM) and Megatron Bridge training with mock or real data.

When should I use Nemo Mbridge Mlm Bridge Training?

Nemo Mbridge Mlm Bridge Training fits situations like: AI & LLM Engineering work in your project.

How do I install Nemo Mbridge Mlm Bridge Training in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-mlm-bridge-training -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-mlm-bridge-training in NVIDIA/skills) into .claude/skills/nemo-mbridge-mlm-bridge-training in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Mlm Bridge Training in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-mlm-bridge-training -a codex`. Or copy the skill folder (skills/nemo-mbridge-mlm-bridge-training in NVIDIA/skills) into .agents/skills/nemo-mbridge-mlm-bridge-training in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Mlm Bridge Training in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-mlm-bridge-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-mlm-bridge-training, .gemini/skills/nemo-mbridge-mlm-bridge-training, .github/skills/nemo-mbridge-mlm-bridge-training and .opencode/skills/nemo-mbridge-mlm-bridge-training in your project.

What does Nemo Mbridge Mlm Bridge Training need to run?

Going by SKILL.md and its folder, Nemo Mbridge Mlm Bridge Training needs the command-line tools its instructions call (uv and git). Our summary lists: Python 3.

Does Nemo Mbridge Mlm Bridge Training access the network?

SKILL.md contains no URLs. Its commands use uv and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Mlm Bridge Training safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Mlm Bridge Training use?

Nemo Mbridge Mlm Bridge Training is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Mlm Bridge Training use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Mlm Bridge Training?

Skills that share tags, products or a category with Nemo Mbridge Mlm Bridge Training: Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars), Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars) and Yolo Detection 2026 (SharpAI/DeepCamera, 3.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Mlm Bridge Training?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.