Official agent skill

Nemo Mbridge Perf Moe Dispatcher Selection

by NVIDIA in NVIDIA/skills

Choose the right MoE token dispatcher (alltoall, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage.

OfficialApache-2.0Auto-check passed

Install Nemo Mbridge Perf Moe Dispatcher Selection

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-dispatcher-selection -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-dispatcher-selection --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-dispatcher-selection .claude/skills/nemo-mbridge-perf-moe-dispatcher-selection && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-moe-dispatcher-selection
GitHub stars
3.6k
Token cost
~2k tokens
SKILL.md length
1,002 words
Files
6
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Choose the right MoE token dispatcher (alltoall, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage.

  • Works in 3 steps: alltoall for initial bring-up → DeepEP if you want a familiar tuned path → HybridEP for the strongest steady-state…
  • SKILL.md covers Quick Decision, Model-Family Patterns, Rounded Evidence Summary and Tuning Parameters, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Nemo Mbridge Perf Moe Dispatcher Selection is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Choose the right MoE token dispatcher (alltoall, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage. Summarizes patterns from DSV3, Qwen3, Qwen3-Next, and VLM bring-up work.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It works with Qwen. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

Example prompts

  • “/nemo-mbridge-perf-moe-dispatcher-selection”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. alltoall for initial bring-up
  2. DeepEP if you want a familiar tuned path
  3. HybridEP for the strongest steady-state result on GB200

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Moe Dispatcher Selection loads about 2k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 1,002 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 1,002 words, ~1,988 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-moe-dispatcher-selection/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-moe-dispatcher-selection
description
Choose the right MoE token dispatcher (`alltoall`, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage. Summarizes patterns from DSV3, Qwen3, Qwen3-Next, and VLM bring-up work.
license
Apache-2.0
when_to_use
Choosing a MoE token dispatcher, or tracing a MoE regression or crash to a dispatcher config change; 'which dispatcher', 'alltoall vs DeepEP', 'HybridEP'…

MoE Dispatcher Selection Guide

Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-dispatcher-selection/card.yaml

Quick Decision

By hardware
HardwareFirst choiceWhy
H100DeepEP, if the runtime package is installedStrong default for cross-node EP on Hopper
B200DeepEP, if the runtime package is installedGood first choice unless a platform-specific HybridEP path is available
GB200 / GB300 NVL72HybridEP, if the runtime package is installedBest fit for NVLink-domain-aware dispatch and lower memory pressure
Unknown or first bring-upalltoallEasiest path for correctness and debugging
By EP degree
EP sizeGuidance
Small EPDispatcher choice is usually second-order; start with alltoall or DeepEP
Medium EPDeepEP often becomes worthwhile
Large EPHybridEP is usually the best target on NVL72 systems

Model-Family Patterns

WorkloadCommon best pathNotes
DSV3 at large scaleHybridEP on GB200 or GB300, DeepEP on H100Dispatcher choice matters more as EP and PP both grow
Qwen3 235BDeepEP on H100, HybridEP on GB200HybridEP usually wins on GB200 and often uses less memory
Qwen3 30BDeepEPSmaller models still benefit, but the absolute gap is smaller
Qwen3-NextClose race in BF16, HybridEP stronger in FP8 or memory-tight runsGood reminder to test, not assume
MoE VLMsStart simple, then test HybridEP on GB200-class systemsVision workloads are sensitive to both memory and host overhead

Rounded Evidence Summary

Backend availability gate

Do not interpret a dispatcher timing until the container has proven that the selected backend package is available. --moe_flex_dispatcher_backend None selects the standard alltoall dispatcher, while deepep and hybridep select moe_token_dispatcher_type="flex" and then require their corresponding runtime packages at model construction time. If DeepEP or HybridEP is missing, record the import failure as an environment limitation and treat alltoall as the only measured correctness fallback for that run.

Qwen3 30B A3B on H100

A short 2026-05-17 H100 smoke run used Qwen3 30B A3B BF16, 16 GPUs, EP=16, the recipe's Transformer Engine CUDA graph scopes (moe_router, moe_preprocess), and model.moe_permute_fusion=false due to a Triton JIT compatibility issue in the run container. The alltoall fallback completed five steps with 45.65 s mean step time after warmup, 132.9 mean TFLOP/s/GPU after warmup, final loss 11.44050, and 61.351 GB peak max allocated memory. DeepEP and HybridEP selected the requested flex backend in the dumped configs but failed before the first iteration because the packages were not installed. This confirms the availability gate; it is not a throughput ranking for flex dispatchers on H100.

DSV3 on GB200 or GB300

The broad trend is more important than any single row in the tracker:

  • plain alltoall is usually the conservative baseline
  • DeepEP improves that baseline once EP communication becomes visible
  • HybridEP adds another step up on NVL72 systems, especially after CUDA graphs, routing improvements, and CPU-side cleanup are already in place

In practice, the stack often moves from roughly "low-teens MFU" territory with an untuned baseline into "high-teens to low-20s MFU" territory after the full dispatcher and kernel stack is tuned.

Qwen3 235B on GB200

For Qwen3 235B, the practical ordering is usually:

  1. alltoall for initial bring-up
  2. DeepEP if you want a familiar tuned path
  3. HybridEP for the strongest steady-state result on GB200

HybridEP is usually modestly faster than alltoall on this workload and often has noticeably better memory headroom.

Qwen3-Next on GB200

This family is a good reminder that dispatcher wins are workload-dependent:

  • in BF16, alltoall and HybridEP can be close
  • in FP8 or memory-constrained settings, HybridEP tends to look better
  • pipeline layout and grouped-GEMM changes can matter almost as much as the dispatcher itself

Tuning Parameters

Show full SKILL.md (420 more words)Show less
DeepEP

DeepEP is selected by setting moe_token_dispatcher_type="flex" and moe_flex_dispatcher_backend="deepep".

bash
--moe-deepep-num-sms 20

Tune the SM count allocated to DeepEP communication kernels (default 20). The optimal value depends on the workload and EP degree. First confirm the DeepEP package imports in the target container; a missing package fails during model construction, before any dispatcher timing is available.

HybridEP

HybridEP is selected by setting moe_token_dispatcher_type="flex" and moe_flex_dispatcher_backend="hybridep".

bash
--moe-hybridep-num-sms 16

Tune the SM count allocated to HybridEP communication (default 16). The performance harness uses 32 for HybridEP workloads. Sweep between 16 and 32 for the target hardware. Set NUM_OF_HYBRID_EP_RANKS_PER_NVLINK_DOMAIN to match the NVLink domain size of the deployment. If it does not match the actual topology, performance and sometimes correctness will suffer. First confirm the HybridEP package imports in the target container; a missing package fails during model construction, before any dispatcher timing is available.

Routing mode
bash
--moe-router-force-load-balancing

For performance benchmarking, force-balance routing is the safer default. It usually outperforms dropless routing in large-scale benchmarks and makes results more comparable across dispatcher backends.

Key Interactions

FeatureInteraction
CUDA graphsBest paired with attn moe_router moe_preprocess on dropless MoE
EP overlapHelps when dispatcher time is still visible after backend tuning
FP8Often increases the relative importance of communication and host overhead
CPU affinityCan matter as much as dispatcher choice on GB200 or GB300
Pipeline layoutPoor PP or VPP layout can erase dispatcher gains

When To Use Each

alltoall
  • first correctness bring-up
  • small EP configurations
  • debugging communication regressions
DeepEP
  • Hopper or B200 deployments
  • cross-node EP is clearly visible in profiles
  • you want a mature intermediate step before testing HybridEP
HybridEP
  • GB200 or GB300 NVL72 systems
  • large EP degrees
  • memory headroom matters in addition to throughput

Pitfalls

  1. Do not compare dispatchers on different stacks: container, routing mode, PP layout, and CUDA-graph scope can move the result as much as the dispatcher.

  2. HybridEP is topology-sensitive: it is not a universal win outside the hardware it was designed for.

  3. Both dispatchers need SM tuning: default moe_deepep_num_sms (20) and moe_hybridep_num_sms (16) are reasonable starting points but rarely optimal.

  4. Force-balance and dropless are not interchangeable baselines: keep the routing mode fixed when comparing dispatcher backends.

  5. Memory and throughput can trade off differently by model: Qwen3-style runs may show a smaller speed delta than DSV3, but still justify HybridEP for memory headroom.

  6. Backend import failures are not performance data: if DeepEP or HybridEP is missing from the container, do not compare its failed job against a completed alltoall job. Fix the environment first, then rerun the same stack.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-dispatcher-selection of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Nemo Mbridge Perf Moe Dispatcher Selection next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Moe Dispatcher Selection compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Moe Dispatcher Selection this skillNVIDIA/skills3.6k—~2kAutomated safety check: PassApache-2.0
AI Image Generation and Editingzhayujie/CowAgent47k—~1.3kAutomated safety check: PassMIT
Agent Feature ReproductionQwenLM/qwen-code28k—~1.5kAutomated safety check: PassApache-2.0
Qwen Code E2E TestingQwenLM/qwen-code28k—~2.1kAutomated safety check: PassApache-2.0
Game Asset Generatorhtdt/godogen7.1k—~2.8kAutomated safety check: PassMIT
Fix Art IssuesOpenPipe/ART11k—~840Automated safety check: NotesApache-2.0

Similar skills

  • Generates or edits images from text prompts through a Python script that picks an image backend based on which API keys are configured.

    47k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Reproduces a feature from Codex or Claude Code in Qwen Code by running the reference agent under capture, reading the traces, then implementing matching behavior.

    28k GitHub stars~1.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Generates game art from text prompts: PNG images, GLB 3D models, rigged characters, animations and sprites, with background removal.

    7.1k GitHub stars~2.8k tokensUpdated 9 days ago
    Game DevelopmentAuto-check passed
  • Fix Art Issues

    OpenPipe/ART

    Fix a GitHub issue on OpenPipe/ART and open a PR. An agent skill from OpenPipe/ART.

    11k GitHub stars~840 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Weavebench Cua Reproduce

    AMAP-ML/LongHorizon-Harness

    Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout.

    1.7k GitHub stars~1.6k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Works with

Questions about Nemo Mbridge Perf Moe Dispatcher Selection

What does Nemo Mbridge Perf Moe Dispatcher Selection do?

Choose the right MoE token dispatcher (alltoall, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage. Nemo Mbridge Perf Moe Dispatcher Selection is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Choose the right MoE token dispatcher (alltoall, DeepEP, or HybridEP) for the hardware, EP degree, and optimization stage.

How do I install Nemo Mbridge Perf Moe Dispatcher Selection in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-dispatcher-selection -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-dispatcher-selection in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-dispatcher-selection in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Moe Dispatcher Selection in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-dispatcher-selection -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-dispatcher-selection in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-dispatcher-selection in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Moe Dispatcher Selection in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-dispatcher-selection -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-dispatcher-selection, .gemini/skills/nemo-mbridge-perf-moe-dispatcher-selection, .github/skills/nemo-mbridge-perf-moe-dispatcher-selection and .opencode/skills/nemo-mbridge-perf-moe-dispatcher-selection in your project.

What does Nemo Mbridge Perf Moe Dispatcher Selection need to run?

SKILL.md names no scripts, command-line tools or credentials: Nemo Mbridge Perf Moe Dispatcher Selection is instructions for the agent only.

Does Nemo Mbridge Perf Moe Dispatcher Selection access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Moe Dispatcher Selection safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Moe Dispatcher Selection use?

Nemo Mbridge Perf Moe Dispatcher Selection is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Moe Dispatcher Selection use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Moe Dispatcher Selection?

Skills that share tags, products or a category with Nemo Mbridge Perf Moe Dispatcher Selection: AI Image Generation and Editing (zhayujie/CowAgent, 47k stars), Agent Feature Reproduction (QwenLM/qwen-code, 28k stars), Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars) and Game Asset Generator (htdt/godogen, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Moe Dispatcher Selection?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.