Official agent skill

Nemo Mbridge Perf Moe Comm Overlap

by NVIDIA in NVIDIA/skills

MoE expert-parallel communication overlap in Megatron Bridge.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Perf Moe Comm Overlap

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-comm-overlap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-comm-overlap .claude/skills/nemo-mbridge-perf-moe-comm-overlap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-moe-comm-overlap
GitHub stars
3.6k
Token cost
~1.9k tokens
SKILL.md length
771 words
Files
6
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

MoE expert-parallel communication overlap in Megatron Bridge.

  • Works in 6 steps: Shared expert overlap conflict:… → PP without VPP: MoE overlap requires VPP… → Flex != backend flag:… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Quick Decision, Enablement, Recompute And CUDA Graph… and Measured Evidence, plus 3 more sections
  • Calls uv

What it does

Nemo Mbridge Perf Moe Comm Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. MoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/nemo-mbridge-perf-moe-comm-overlap”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Shared expert overlap conflict: moe_shared_expert_overlap and
  2. PP without VPP: MoE overlap requires VPP when pipeline parallelism is
  3. Flex != backend flag: moe_flex_dispatcher_backend="deepep" alone
  4. Conservative recipe defaults: Most public recipes leave MoE overlap
  5. Performance gains are workload-dependent: overlap helps most when dispatch
  6. Summed kernel time is not wall time: concurrent kernels can run longer

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Moe Comm Overlap loads about 1.9k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 771 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 771 words, ~1,873 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-moe-comm-overlap/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-moe-comm-overlap
description
MoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
license
Apache-2.0
when_to_use
Tuning MoE communication overlap, or tracing a MoE throughput regression to a comm-overlap config change; 'overlap_moe_expert_parallel_comm', 'MoE dispatch…

MoE Communication Overlap

For the higher-level overview, see:

  • @docs/training/communication-overlap.md
  • @skills/nemo-mbridge-perf-moe-comm-overlap/card.yaml

Quick Decision

Use MoE communication overlap when:

  • EP > 1
  • token dispatch or combine time is visible in the profile
  • the run is already correct and you are now tuning throughput

Avoid turning it on as an early bring-up step. It is easier to validate after the dispatcher, routing mode, and recompute plan are already stable.

Enablement

python
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True

# Optional: delayed wgrad for additional overlap
cfg.comm_overlap.delay_wgrad_compute = True

# IMPORTANT: disable shared expert overlap when using dispatch overlap
cfg.model.moe_shared_expert_overlap = False
Prerequisites
  • expert_model_parallel_size > 1
  • num_moe_experts > 1
  • moe_token_dispatcher_type must be "alltoall" or "flex"
  • Precision: BF16 or FP16
  • If PP is used, VPP (virtual_pipeline_model_parallel_size) must be set (non-None)
Flex dispatcher activation

Setting moe_flex_dispatcher_backend alone does not activate flex dispatch. You must also set moe_token_dispatcher_type = "flex".

Recompute And CUDA Graph Interaction

  • Full recompute is not a good companion for the overlap path.
  • delay_wgrad_compute adds further constraints if CUDA-graph scopes include attention or MoE-router work.
  • In practice, selective recompute is the safer pairing when overlap is enabled.

Measured Evidence

HybridEP production-shape validation

A 2026-07-25 controlled Qwen3 30B-A3B pretraining comparison used 16 H100 GPUs, BF16, sequence length 4096, TP=1, PP=1, CP=1, EP=16, MBS=1, GBS=1024, forced-balanced routing, HybridEP, and Transformer Engine CUDA-graph scopes moe_router and moe_preprocess. The only performance change was plain EP overlap; delayed wgrad stayed disabled.

CaseSteady windowStep timeModel TFLOPS/GPU
EP overlap offiterations 5-2024.7138s244.039
EP overlap on, search runiterations 5-2021.0725s286.208
EP overlap on, independent validationiterations 41-5020.9920s287.305

The independent result reduced step time by 15.059% and increased throughput by 17.729% over the reproduced baseline. Loss remained finite, no iterations were skipped or NaN, and rank-0 peak allocated memory was 62.166 GiB.

A same-method rank-0 Nsight Systems comparison captured 463,348 kernels in each case:

Profile metricOverlap offOverlap on
Communication concurrent with GEMM/attention9.079ms3,958.997ms
Communication time hidden by compute0.11%36.55%
GPU-active interval union22.821s21.221s
HybridEP dispatch-with-permute NVTX4.253s1.767s
HybridEP metadata-preprocess NVTX3.109s0.670s

This is direct evidence that the gain came from hiding exposed HybridEP dispatch/combine work, not from changing the dispatcher, routing, graph scopes, batch shape, or parallel layout.

Correctness-first alltoall smoke

A 2026-05-18 current-main H100 x16 smoke on Qwen3 30B-A3B mock pretraining used EP=16, alltoall, global batch size 1024, CUDA graphs disabled, and moe_permute_fusion=false because the PyTorch 25.11 / TE / Triton stack failed in Transformer Engine fused permutation in prior bring-up.

Results were directional rather than release-grade:

  • no EP overlap: 41.25s steady-state mean over iterations 3-8
  • EP overlap: 31.31s steady-state mean over iterations 3-8
  • EP overlap plus delay_wgrad_compute: 31.20s steady-state mean over iterations 3-8

Treat this as evidence that EP overlap can help an inter-node alltoall MoE shape when communication is exposed. It is not proof that delayed wgrad is a separate win, and it does not validate the fused permutation path. An earlier 2026-05-16 short smoke on the same shape showed the same pattern.

Show full SKILL.md (309 more words)Show less

Code Anchors

  • Overlap validation: src/megatron/bridge/training/comm_overlap.py
  • Flex dispatcher backend: src/megatron/bridge/training/flex_dispatcher_backend.py
  • Config: src/megatron/bridge/training/config.py
  • Unit tests: tests/unit_tests/training/test_comm_overlap.py
  • DeepEP tests: tests/unit_tests/training/test_deepep.py

Pitfalls

  1. Shared expert overlap conflict: moe_shared_expert_overlap and overlap_moe_expert_parallel_comm can conflict. Disable shared expert overlap when using the dispatch overlap path.

  2. PP without VPP: MoE overlap requires VPP when pipeline parallelism is active. Without it, the overlap scheduling cannot interleave correctly.

  3. Flex != backend flag: moe_flex_dispatcher_backend="deepep" alone does nothing if moe_token_dispatcher_type is still "alltoall".

  4. Conservative recipe defaults: Most public recipes leave MoE overlap disabled. You need to explicitly enable it via overrides.

  5. Performance gains are workload-dependent: overlap helps most when dispatch communication is already a visible slice of step time. It is not guaranteed to help every small or lightly loaded EP run.

  6. Summed kernel time is not wall time: concurrent kernels can run longer because they contend for SMs or bandwidth, so overlap may increase summed per-stream kernel duration while reducing the exposed interval union and end-to-end step time.

Verification

Look for overlap-related log messages during initialization. The comm overlap validation in comm_overlap.py will raise if prerequisites are not met, so a clean startup confirms the feature is active.

For a short performance-harness smoke, keep the command shape explicit and vary only one overlap knob at a time:

bash
uv run python scripts/performance/run_script.py \
  -m qwen \
  -mr qwen3_30b_a3b \
  --task pretrain \
  -g h100 \
  -c bf16 \
  -ng 16 \
  -gn 8 \
  --max_steps 8 \
  --cuda_graph_impl none \
  --moe_flex_dispatcher_backend None \
  --moe_a2a_overlap false \
  --tokenizer_type NullTokenizer \
  comm_overlap.overlap_moe_expert_parallel_comm=true \
  comm_overlap.delay_wgrad_compute=false \
  model.moe_shared_expert_overlap=false

If fused MoE permutation fails during bring-up, add model.moe_permute_fusion=false to separate overlap timing from runtime-stack validation, then retest with the matched production container.

For performance validation, use an unprofiled steady window as the acceptance metric. Use a matched Nsight A/B to establish causality:

  1. Keep dispatcher, routing, CUDA graphs, batch shape, parallelism, and runtime fixed.
  2. Toggle only overlap_moe_expert_parallel_comm; keep delay_wgrad_compute=false for the first isolation.
  3. Compare communication and compute interval unions and their intersection, not only summed kernel durations.
  4. Report steady step time, model TFLOPS/GPU, loss finiteness, skipped/NaN iterations, and peak allocated memory.

Last signature refresh: 2026-08-03.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-comm-overlap of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Nemo Mbridge Perf Moe Comm Overlap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Moe Comm Overlap compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Moe Comm Overlap this skillNVIDIA/skills3.6k—~1.9kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    188 GitHub stars~547 tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Nemo Mbridge Perf Moe Comm Overlap

What does Nemo Mbridge Perf Moe Comm Overlap do?

MoE expert-parallel communication overlap in Megatron Bridge. Nemo Mbridge Perf Moe Comm Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. MoE expert-parallel communication overlap in Megatron Bridge.

When should I use Nemo Mbridge Perf Moe Comm Overlap?

Nemo Mbridge Perf Moe Comm Overlap fits situations like: AI & LLM Engineering work in your project.

How do I install Nemo Mbridge Perf Moe Comm Overlap in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-comm-overlap in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-comm-overlap in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Moe Comm Overlap in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-comm-overlap in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-comm-overlap in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Moe Comm Overlap in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-comm-overlap, .gemini/skills/nemo-mbridge-perf-moe-comm-overlap, .github/skills/nemo-mbridge-perf-moe-comm-overlap and .opencode/skills/nemo-mbridge-perf-moe-comm-overlap in your project.

What does Nemo Mbridge Perf Moe Comm Overlap need to run?

Going by SKILL.md and its folder, Nemo Mbridge Perf Moe Comm Overlap needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Perf Moe Comm Overlap access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Moe Comm Overlap safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Moe Comm Overlap use?

Nemo Mbridge Perf Moe Comm Overlap is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Moe Comm Overlap use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Moe Comm Overlap?

Skills that share tags, products or a category with Nemo Mbridge Perf Moe Comm Overlap: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Moe Comm Overlap?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.