Official agent skill

Nemo Mbridge Perf Expert Parallel Overlap

by NVIDIA in NVIDIA/skills

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Perf Expert Parallel Overlap

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-expert-parallel-overlap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-expert-parallel-overlap .claude/skills/nemo-mbridge-perf-expert-parallel-overlap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-expert-parallel-overlap
GitHub stars
3.5k
Token cost
~3.5k tokens
SKILL.md length
1,186 words
Files
6
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.

  • Works in 3 steps: Confirm no assertion errors during… → Confirm overlap_moe_expert_parallel_comm… → If using flex dispatcher, confirm…
  • AI & LLM Engineering work in your project
  • SKILL.md covers References, What It Is, Quick Decision and Correctness-First alltoall…, plus 9 more sections
  • Calls uv

What it does

Nemo Mbridge Perf Expert Parallel Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/nemo-mbridge-perf-expert-parallel-overlap”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Confirm no assertion errors during CommOverlapConfig finalization
  2. Confirm overlap_moe_expert_parallel_comm appears as True in the logged
  3. If using flex dispatcher, confirm moe_token_dispatcher_type = "flex" and

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Expert Parallel Overlap loads about 3.5k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 1,186 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~61
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,186 words, ~3,536 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-expert-parallel-overlap/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-expert-parallel-overlap
description
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as DeepEP and HybridEP.
license
Apache-2.0
when_to_use
Enabling EP overlap to hide dispatch/combine latency, or tracing a throughput regression to an EP overlap config change; 'overlap_moe_expert_parallel_comm'…

MoE Expert-Parallel Overlap Skill

References

  • Stable docs: @docs/training/communication-overlap.md
  • Structured metadata: @skills/nemo-mbridge-perf-expert-parallel-overlap/card.yaml

What It Is

Expert-parallel (EP) overlap hides the cost of token dispatch/combine all-to-all communication by running it concurrently with expert FFN compute. Optionally, delayed expert weight-gradient computation (delay_wgrad_compute) provides additional overlap by deferring wgrad to overlap with the next layer's forward.

Bridge supports two dispatcher paths:

DispatcherBackendWhen to use
alltoallStandard MoE all-to-allDefault, broadest compatibility
flexDeepEP or HybridEPHigher overlap on Ampere/Hopper/Blackwell

Quick Decision

Use EP overlap when:

  • the model is MoE with EP > 1
  • expert dispatch/combine communication is a meaningful part of step time
  • you have memory headroom and are tuning for throughput

Prefer:

  • alltoall dispatcher for the first rollout (broader compatibility)
  • flex + DeepEP/HybridEP when running on supported GPUs and seeking additional gains

Avoid EP overlap when:

  • full activation recompute is enabled
  • moe_shared_expert_overlap is enabled
  • the run is still being brought up for correctness
  • PyTorch < 2.6.0

Expected outcome:

  • if all-to-all dispatch is a clear profile bottleneck, overlap can produce a modest to meaningful speedup
  • if the run is tiny, communication-light, or dominated by another wall, the gain may be negligible

Correctness-First alltoall Benchmark

For the plain EP-overlap isolation benchmark, keep flex dispatch and delayed wgrad disabled. The measured shape was Qwen3 MoE 30B-A3B SFT on 16 H100 GPUs: EP=16, alltoall, BF16, global batch size 1024, CUDA graphs disabled, moe_permute_fusion=false, measured over iterations 3-8.

Use these overrides for the plain-overlap case:

bash
--cuda_graph_impl none \
--moe_flex_dispatcher_backend None \
--moe_a2a_overlap false \
comm_overlap.overlap_moe_expert_parallel_comm=true \
comm_overlap.delay_wgrad_compute=false \
model.moe_shared_expert_overlap=false

Do not use --moe_a2a_overlap true for this isolation test: the performance harness helper enables both overlap_moe_expert_parallel_comm and delay_wgrad_compute, so it does not isolate plain EP overlap.

Steady-window timing from that benchmark:

CaseSteady meanRelative
no EP overlap41.25s1.000x
EP overlap31.31s1.317x
EP overlap plus delay_wgrad_compute31.20s1.322x

This is evidence for enabling plain EP overlap on this inter-node all-to-all shape. It does not show a meaningful independent win from delayed wgrad, and it does not validate fused MoE permutation because that path was disabled for the runtime stack.

HybridEP Production-Shape Benchmark

A 2026-07-25 controlled Qwen3 30B-A3B pretraining comparison validated plain EP overlap with the production HybridEP path:

text
Hardware: 16×H100
Precision: BF16
Sequence: 4096
Parallelism: TP1 / PP1 / CP1 / EP16
Batch: MBS1 / GBS1024
Routing: force balance
Dispatcher: flex + HybridEP
CUDA graph: Transformer Engine scopes moe_router + moe_preprocess
Delayed wgrad: disabled
CaseSteady windowStep timeModel TFLOPS/GPU
overlap offiterations 5-2024.7138s244.039
overlap on, search runiterations 5-2021.0725s286.208
overlap on, independent validationiterations 41-5020.9920s287.305

The independent run reduced step time by 15.059% and raised throughput by 17.729% over the reproduced baseline. Loss was finite, skipped and NaN iterations remained zero, and rank-0 peak allocated memory was 62.166 GiB.

A matched Nsight Systems comparison captured the same 463,348 rank-0 kernels per case. Enabling overlap increased communication concurrent with GEMM and attention from 9.079ms (0.11% of communication time) to 3,958.997ms (36.55%). GPU-active interval union fell from 22.821s to 21.221s.

Use this as evidence for the mechanism, not as a universal speedup promise. The dispatcher, graph scopes, routing, parallelism, batch shape, and runtime were held fixed while only plain EP overlap changed.

Enablement

alltoall dispatcher
python
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.moe_shared_expert_overlap = False

cfg.model.expert_model_parallel_size = 8
cfg.model.num_moe_experts = 64
cfg.model.moe_token_dispatcher_type = "alltoall"
cfg.model.bf16 = True
cfg.model.fp16 = False

Enable delay_wgrad_compute=True only after the plain overlap path is known to work and its extra compatibility constraints have been checked.

flex dispatcher (DeepEP or HybridEP)
python
from megatron.bridge.training.flex_dispatcher_backend import apply_flex_dispatcher_backend

cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.moe_shared_expert_overlap = False

apply_flex_dispatcher_backend(cfg.model, moe_flex_dispatcher_backend="deepep")
# or: apply_flex_dispatcher_backend(cfg.model, moe_flex_dispatcher_backend="hybridep")

Benchmark plain EP overlap first. Enable delay_wgrad_compute=True only as a separate follow-up A/B after its CUDA-graph and TE compatibility constraints are satisfied.

Compatibility And Constraints

  • expert_model_parallel_size > 1
  • num_moe_experts > 1
  • moe_token_dispatcher_type must be "alltoall" or "flex"
  • moe_shared_expert_overlap = False
  • Base precision is BF16 or FP16
  • PyTorch >= 2.6.0
  • If PP > 1, virtual_pipeline_model_parallel_size must be set
  • recompute_granularity != "full", recompute_method = None, recompute_num_layers = None
  • mtp_num_layers must be None or 1
  • delay_wgrad_compute requires overlap_moe_expert_parallel_comm as a prerequisite
  • delay_wgrad_compute with overlap_grad_reduce requires TE >= 2.7.0
  • delay_wgrad_compute with gradient_accumulation_fusion requires TE >= 2.7.0
  • CUDA graph attn scope + delay_wgrad_compute requires TE >= 2.12.0, gradient_accumulation_fusion = True, and no attention bias
  • DeepEP: Ampere, Hopper, B200, B300 GPUs only
  • HybridEP: Ampere, Hopper, B200, B300, GB200/GB300 with NVL72

Minimal Working Config

python
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.expert_model_parallel_size = 4
cfg.model.num_moe_experts = 64
cfg.model.moe_token_dispatcher_type = "alltoall"
cfg.model.moe_shared_expert_overlap = False
cfg.model.bf16 = True

Use this as the correctness-first starting point. Add delayed wgrad, flex dispatch, and CUDA-graph interactions only after the plain overlap path is known to work.

Minimal Runnable Command

Performance harness example inside a Slurm allocation. Keep the model, parallelism, dispatcher, and runtime fixed, and vary only the two overlap overrides:

bash
uv run python scripts/performance/run_script.py \
  -m qwen \
  -mr qwen3_30b_a3b \
  --task pretrain \
  -g h100 \
  -c bf16 \
  -ng 16 \
  -gn 8 \
  --max_steps 8 \
  --cuda_graph_impl none \
  --moe_flex_dispatcher_backend None \
  --moe_a2a_overlap false \
  --tokenizer_type NullTokenizer \
  comm_overlap.overlap_moe_expert_parallel_comm=true \
  comm_overlap.delay_wgrad_compute=false \
  model.moe_shared_expert_overlap=false

Do not use --moe_a2a_overlap true when separating plain EP overlap from delayed wgrad: the performance harness helper enables both overlap_moe_expert_parallel_comm and delay_wgrad_compute.

Unit test verification:

bash
uv run python -m pytest \
  tests/unit_tests/training/test_comm_overlap.py -k "moe" \
  tests/unit_tests/training/test_deepep.py -q

Verification

Unit tests
bash
uv run python -m pytest \
  tests/unit_tests/training/test_comm_overlap.py \
  tests/unit_tests/training/test_deepep.py -q
Show full SKILL.md (486 more words)Show less
Log checks

After a successful run with EP overlap:

  1. Confirm no assertion errors during CommOverlapConfig finalization
  2. Confirm overlap_moe_expert_parallel_comm appears as True in the logged config
  3. If using flex dispatcher, confirm moe_token_dispatcher_type = "flex" and the correct backend in logs
Success criteria
  • Config validation passes for the selected dispatcher and overlap settings
  • Training runs complete without hangs or assertion failures
  • Throughput improves or at least does not regress for the target workload
  • Loss trajectory matches baseline (overlap should not affect convergence)
Profile interpretation

Use an unprofiled steady window for the throughput acceptance result. Use a matched profile to explain the mechanism:

  1. Keep the dispatcher, routing, graph scopes, batch shape, parallel layout, and runtime fixed.
  2. Capture the same rank and steady iteration while toggling only plain EP overlap.
  3. Build interval unions for communication and compute kernels, then measure their intersection.
  4. Do not use summed kernel duration as wall time. Concurrent kernels can run longer under SM or bandwidth contention even when exposed time decreases.
  5. Corroborate interval results with dispatch/combine NVTX ranges, final step time, loss finiteness, skipped/NaN counts, and peak memory.

Code Anchors

Bridge overlap validation
470505srcmegatronbri
if self.user_comm_overlap_cfg.overlap_moe_expert_parallel_comm is True:
    assert model_cfg.expert_model_parallel_size > 1, ...
    assert model_cfg.num_moe_experts > 1, ...
    assert model_cfg.moe_token_dispatcher_type in ["alltoall", "flex"], ...
    assert model_cfg.bf16 or model_cfg.fp16, ...
    assert is_torch_min_version("2.6.0"), ...
    # ... PP + VPP check, recompute checks, shared_expert_overlap check ...
Delayed wgrad validation
507557srcmegatronbri
if self.user_comm_overlap_cfg.delay_wgrad_compute is True:
    # TE version checks for overlap_grad_reduce and gradient_accumulation_fusion
    # CUDA graph scope validations for delayed wgrad
    assert overlap_moe_expert_parallel_comm, ...
Flex-dispatcher activation
2772srcmegatronbridg
def apply_flex_dispatcher_backend(...):
    # GPU architecture check for DeepEP / HybridEP
    model_config.moe_token_dispatcher_type = "flex"
    model_config.moe_flex_dispatcher_backend = moe_flex_dispatcher_backend
    model_config.moe_shared_expert_overlap = False
Perf harness override
149156scriptsperform
def _set_moe_a2a_overlap_overrides(recipe, moe_a2a_overlap=False):
    if moe_a2a_overlap:
        recipe.comm_overlap.overlap_moe_expert_parallel_comm = True
        recipe.comm_overlap.delay_wgrad_compute = True
        recipe.model.moe_shared_expert_overlap = False
Tests
FileCoverage
tests/unit_tests/training/test_comm_overlap.pyEP overlap validation, delayed wgrad, CUDA graph + wgrad interaction
tests/unit_tests/training/test_deepep.pyDeepEP/HybridEP helper activation and GPU gating

Failure Diagnosis

SymptomLikely CauseHow To ConfirmFix
assert expert_model_parallel_size > 1EP not configuredCheck expert_model_parallel_sizeSet EP > 1
assert moe_token_dispatcher_typeWrong dispatcherCheck dispatcher typeUse "alltoall" or "flex"
assert on BF16/FP16Wrong precisionCheck bf16 and fp16Set bf16 = True
hang during trainingPyTorch < 2.6Check PyTorch versionUpgrade to >= 2.6.0
assert virtual_pipeline_model_parallel_sizePP > 1 without VPPCheck PP and VPP configSet VPP when PP > 1
assert recompute_granularityFull recompute enabledCheck recompute settingsDisable full recompute
assert overlap_moe_expert_parallel_comm requireddelayed wgrad without EP overlapCheck delay_wgrad_compute without overlapEnable EP overlap first
assert gradient_accumulation_fusionCUDA graph + delayed wgradCheck graph scope + wgrad settingsEnable gradient_accumulation_fusion
assert on attention biasCUDA graph attn + delayed wgrad + biasCheck add_bias_linear / add_qkv_biasDisable attention bias
no throughput gain from flex dispatcherapply_flex_dispatcher_backend not calledCheck moe_token_dispatcher_type in logsCall apply_flex_dispatcher_backend(...)
DeepEP/HybridEP silently skippedUnsupported GPUCheck warning logsRun on Ampere/Hopper/Blackwell
summed kernel time increases after overlapExpected concurrency contention or a regressionCompare interval unions, comm/compute intersection, and unprofiled step timeJudge overlap from exposed wall time, not summed per-stream duration

Known Limitations

  • Setting moe_flex_dispatcher_backend alone does not activate flex dispatch — you must call apply_flex_dispatcher_backend(...).
  • Public recipes are often conservative and leave MoE overlap disabled by default.
  • Controlled end-to-end and profile evidence exists for one Qwen3 30B-A3B HybridEP H100 shape; repeat the matched A/B before generalizing it to another model, dispatcher, topology, precision, or batch shape.
  • MoE overlap and shared-expert overlap are mutually exclusive.
  • CUDA graph plus delayed wgrad is a multi-constraint path that requires careful TE version and scope validation.

Last signature refresh: 2026-08-03.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-expert-parallel-overlap of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemo Mbridge Perf Expert Parallel Overlap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Expert Parallel Overlap compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Expert Parallel Overlap this skillNVIDIA/skills3.5k—~3.5kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    170 GitHub stars~547 tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Nemo Mbridge Perf Expert Parallel Overlap

What does Nemo Mbridge Perf Expert Parallel Overlap do?

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. Nemo Mbridge Perf Expert Parallel Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.

When should I use Nemo Mbridge Perf Expert Parallel Overlap?

Nemo Mbridge Perf Expert Parallel Overlap fits situations like: AI & LLM Engineering work in your project.

How do I install Nemo Mbridge Perf Expert Parallel Overlap in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-expert-parallel-overlap in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-expert-parallel-overlap in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Expert Parallel Overlap in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-expert-parallel-overlap in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-expert-parallel-overlap in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Expert Parallel Overlap in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-expert-parallel-overlap, .gemini/skills/nemo-mbridge-perf-expert-parallel-overlap, .github/skills/nemo-mbridge-perf-expert-parallel-overlap and .opencode/skills/nemo-mbridge-perf-expert-parallel-overlap in your project.

What does Nemo Mbridge Perf Expert Parallel Overlap need to run?

Going by SKILL.md and its folder, Nemo Mbridge Perf Expert Parallel Overlap needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Perf Expert Parallel Overlap access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Expert Parallel Overlap safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Expert Parallel Overlap use?

Nemo Mbridge Perf Expert Parallel Overlap is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Expert Parallel Overlap use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Expert Parallel Overlap?

Skills that share tags, products or a category with Nemo Mbridge Perf Expert Parallel Overlap: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Expert Parallel Overlap?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.