Official agent skill

Nemo Mbridge Perf Moe Hardware Configs

by NVIDIA in NVIDIA/skills

Representative, point-in-time MoE training playbooks by hardware and model family.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Perf Moe Hardware Configs

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-hardware-configs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-hardware-configs .claude/skills/nemo-mbridge-perf-moe-hardware-configs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-moe-hardware-configs
GitHub stars
3.6k
Token cost
~2k tokens
SKILL.md length
722 words
Files
6
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Representative, point-in-time MoE training playbooks by hardware and model family.

  • Works in 7 steps: Do not cargo-cult a tracker row: the… → Container quality matters: large… → VPP must be intentional: a bad VPP split… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Quick Platform Playbook, First Answer Checklist, Rounded Performance Bands and Representative Config Families, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Nemo Mbridge Perf Moe Hardware Configs is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Representative, point-in-time MoE training playbooks by hardware and model family. Use them as candidate seeds, then revalidate the exact runtime, semantics, topology, and steady-state throughput.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering. It works with CUDA and Qwen. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/nemo-mbridge-perf-moe-hardware-configs”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Do not cargo-cult a tracker row: the winning config usually depends on
  2. Container quality matters: large regressions can come from the software
  3. VPP must be intentional: a bad VPP split can erase the gain from a better
  4. Compare absolute throughput, not only MFU: MFU can mislead when switching
  5. Force-balance routing is benchmark-only: it can control routing variance,
  6. Do not treat the dispatcher table as a hard platform rule: HybridEP is
  7. Separate screening, causality, and acceptance: short runs reject weak

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Moe Hardware Configs loads about 2k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 722 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 722 words, ~1,984 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-moe-hardware-configs/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-moe-hardware-configs
description
Representative, point-in-time MoE training playbooks by hardware and model family. Use them as candidate seeds, then revalidate the exact runtime, semantics, topology, and steady-state throughput.
license
Apache-2.0
when_to_use
Hardware-specific MoE playbooks or throughput estimates; 'MoE on H100', 'GB200 config', 'expected throughput', 'MoE hardware playbook', 'parallelism for B200'.

MoE Hardware Configuration Reference

Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-hardware-configs/card.yaml

Quick Platform Playbook

These rows are search seeds, not hardware defaults or throughput promises.

PlatformCandidates to screen after alltoall bring-upWhat usually matters most
H100DeepEP or HybridEP, explicit overlap, supported FP8 modescommunication overlap, dispatcher/runtime compatibility, and PP efficiency
B200DeepEP or HybridEP, supported FP8 modes, careful PP layoutcontainer quality and tuned communication settings
GB200HybridEP, then profile-driven graphs and CPU cleanuphost overhead, topology-aware dispatch, memory headroom
GB300HybridEP and the target container's lower-precision/kernel stackthe same system interactions as GB200, with remeasurement required

First Answer Checklist

For hardware playbook questions, answer from these canonical rows before adding throughput caveats:

WorkloadHardwareDispatcherLayout
DSV3H100DeepEPTP=2, EP=64, PP=8, VPP=4
DSV3GB200/GB300HybridEPTP=1, EP=64, PP=4, VPP=4
Qwen3 235BH100alltoall + overlap in the current canonical recipeTP=2, EP=32, PP=8, VPP=4
Qwen3 235BGB200HybridEPTP=1 or 2, EP=32-64, PP=4, VPP=unspecified
Qwen3 30B16×H100HybridEPTP=1, EP=16, PP=1, plain EP overlap

For Qwen3 235B on GB200, explicitly say VPP=unspecified; do not invent or extrapolate VPP=12 unless a measured row provides it. Treat TE-scoped CUDA graph scopes (attn, moe_router, moe_preprocess) as profile-driven candidates, CUDA_DEVICE_MAX_CONNECTIONS selection, PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, NCCL_GRAPH_REGISTER=0, GB200/GB300 CPU-side tuning, and the warning not to cargo-cult tracker rows.

Rounded Performance Bands

These are intentionally rounded so the document stays durable as the tracker moves. Treat them as planning ranges, not exact promises.

Workload familyHardwareTypical bandRepresentative shape
DSV3, large-scaleH100low-to-mid hundreds TFLOPS/GPU, high-teens MFUTP2, EP64, PP8, DeepEP
DSV3, large-scaleB200high-hundreds TFLOPS/GPU, mid-teens MFUTP1, EP32, PP8, DeepEP
DSV3, large-scaleGB200around 1K TFLOPS/GPU, low-20s MFUTP1, EP64, PP4, HybridEP
DSV3, large-scaleGB300above the GB200 band, often mid-20s MFUTP1, EP64, PP4, HybridEP
Qwen3 235BH100historical low-300s snapshots; remeasure the current recipeTP2, EP32, PP8; current recipe uses alltoall + overlap
Qwen3 235BGB200high-hundreds TFLOPS/GPU in tuned runsTP1 or TP2, EP32-64, PP4, HybridEP
Qwen3 30BH100about 300 TFLOPS/GPU on the validated 16-GPU shapeTP1, EP16, PP1, HybridEP + EP overlap
Qwen3-Next 80BGB200low-300s TFLOPS/GPU in BF16-class runsTP1, EP32, PP2, HybridEP

Representative Config Families

DSV3 on H100
text
Dispatcher: DeepEP
TP=2  EP=64  PP=8  VPP=4
Routing: force balance
Recompute: light-to-moderate selective recompute
Priority: overlap communication and keep PP efficient
DSV3 on B200
text
Dispatcher: DeepEP
TP=1  EP=32  PP=8  VPP=2 or similar
Precision: MXFP8-class
Recompute: selective recompute around MLA up-projection and MLP-side modules
Priority: container quality, PP layout, and DeepEP SMS tuning
DSV3 on GB200 or GB300
text
Dispatcher: HybridEP
TP=1  EP=64  PP=4  VPP=4
Precision: MXFP8-class
CUDA Graph: attn + moe_router + moe_preprocess
Priority: HybridEP, CPU optimization, and graph-friendly static shapes
Qwen3 235B on H100
text
Dispatcher: alltoall in the current canonical recipe; re-screen flex backends on the target stack
TP=2  EP=32  PP=8  VPP=4
Recompute: none in the current canonical recipe
Priority: communication overlap and router-path cleanup
Qwen3 235B on GB200
text
Dispatcher: HybridEP
TP=1 or 2  EP=32 to 64  PP=4  VPP=unspecified unless measured
CUDA Graph: attn + moe_router + moe_preprocess
Recompute: moe_act, mlp, or norm depending on memory pressure
Priority: balance throughput against memory headroom
Qwen3 30B-A3B on 16 H100
text
Dispatcher: HybridEP
TP=1  EP=16  PP=1  CP=1
Precision: BF16
Sequence: 4096
Batch: MBS1 GBS1024
Routing: force balance
EP overlap: enabled
Delayed wgrad: disabled
CUDA Graph: moe_router + moe_preprocess
HybridEP: permute fusion, 32 SMs, 64-token combine chunks
Measured: 20.14729s/step, 299.352 model TFLOPS/GPU over iterations 41-50
Rank-0 peak allocated memory: 62.166 GiB

The current number is the final multi-knob canonical recipe result. An earlier matched A/B isolated plain EP overlap: 244.039 to 287.305 TFLOPS/GPU, with communication hidden by GEMM/attention increasing from 0.11% to 36.55%. Do not attribute the later 299.352 result entirely to overlap.

Qwen3-Next 80B on GB200
text
Dispatcher: HybridEP
TP=1  EP=32  PP=2  VPP around 4
CUDA Graph: attn + moe_router + moe_preprocess
Priority: pipeline layout and grouped GEMM quality

Cross-Cutting Patterns

Show full SKILL.md (295 more words)Show less
PP layout
  • E = embedding
  • t = transformer
  • m = MTP
  • L = loss
  • | = stage boundary

The biggest platform difference is usually not just the dispatcher. It is the combination of dispatcher, PP shape, and whether VPP keeps each stage balanced.

Recompute strategy
Memory pressureStarting point
lownone or a very narrow selective set
moderatemoe_act, mlp, norm, or similar selective modules
highmodel-specific up-projection plus selective MoE and MLP modules
extreme or long-contextfull recompute only if the selective path still does not fit
Environment variables
bash
CUDA_DEVICE_MAX_CONNECTIONS=1
CUDA_DEVICE_MAX_CONNECTIONS=32   # common when EP overlap and CUDA graphs are combined
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
NCCL_GRAPH_REGISTER=0
CPU-side tuning

On GB200 and GB300, CPU affinity and general host-overhead cleanup can move the needle almost as much as a dispatcher swap. Treat them as first-class tuning work, not as afterthoughts.

Pitfalls

  1. Do not cargo-cult a tracker row: the winning config usually depends on routing mode, container, and PP layout as much as on hardware name.

  2. Container quality matters: large regressions can come from the software stack rather than the model recipe.

  3. VPP must be intentional: a bad VPP split can erase the gain from a better dispatcher.

  4. Compare absolute throughput, not only MFU: MFU can mislead when switching between BF16, FP8, and other precision modes.

  5. Force-balance routing is benchmark-only: it can control routing variance, but it changes semantics. Keep routing fixed within an A/B and validate natural routing separately for training acceptance.

  6. Do not treat the dispatcher table as a hard platform rule: HybridEP is the validated winner for the canonical 16×H100 Qwen3 30B shape, while the current 256×H100 Qwen3 235B recipe uses alltoall. Benchmark backend compatibility and throughput in the production container.

  7. Separate screening, causality, and acceptance: short runs reject weak candidates, matched one-variable A/Bs explain a mechanism, and a 50-step final run validates the complete winner.

Last signature refresh: 2026-08-03.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-hardware-configs of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Nemo Mbridge Perf Moe Hardware Configs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Moe Hardware Configs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Moe Hardware Configs this skillNVIDIA/skills3.6k—~2kAutomated safety check: PassApache-2.0
Add New ModelJakeATX/llamAmpere157—~4.1kAutomated safety check: PassMIT
Debug Cuda Crashsgl-project/sglang37k2 repos~4.9kAutomated safety check: PassApache-2.0
Code ReviewJakeATX/llamAmpere157—~5.6kAutomated safety check: PassMIT
Add Modelsohu-mptc/FlashRec107—~1.7kAutomated safety check: PassApache-2.0
AppJakeATX/llamAmpere157—~146Automated safety check: PassMIT

Similar skills

  • Add New Model

    JakeATX/llamAmpere

    Guided workflow for adding a new model architecture to llama.cpp.

    157 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Debug Cuda Crash

    sgl-project/sglang

    Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

    37k GitHub starsUsed in 2 repos~4.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Code Review

    JakeATX/llamAmpere

    Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.

    157 GitHub stars~5.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Add Model

    sohu-mptc/FlashRec

    给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。

    107 GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • App

    JakeATX/llamAmpere

    Opinionated app components building on top of ./ui primitives

    157 GitHub stars~146 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Market Data

    zhongkaifu/TensorSharp

    Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

    563 GitHub stars~1k tokensUpdated today
    Business, Finance & HRAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Works with

Questions about Nemo Mbridge Perf Moe Hardware Configs

What does Nemo Mbridge Perf Moe Hardware Configs do?

Representative, point-in-time MoE training playbooks by hardware and model family. Nemo Mbridge Perf Moe Hardware Configs is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Representative, point-in-time MoE training playbooks by hardware and model family.

When should I use Nemo Mbridge Perf Moe Hardware Configs?

Nemo Mbridge Perf Moe Hardware Configs fits situations like: AI & LLM Engineering work in your project.

How do I install Nemo Mbridge Perf Moe Hardware Configs in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-hardware-configs in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-hardware-configs in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Moe Hardware Configs in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-hardware-configs in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-hardware-configs in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Moe Hardware Configs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-hardware-configs, .gemini/skills/nemo-mbridge-perf-moe-hardware-configs, .github/skills/nemo-mbridge-perf-moe-hardware-configs and .opencode/skills/nemo-mbridge-perf-moe-hardware-configs in your project.

What does Nemo Mbridge Perf Moe Hardware Configs need to run?

SKILL.md names no scripts, command-line tools or credentials: Nemo Mbridge Perf Moe Hardware Configs is instructions for the agent only.

Does Nemo Mbridge Perf Moe Hardware Configs access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Moe Hardware Configs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Moe Hardware Configs use?

Nemo Mbridge Perf Moe Hardware Configs is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Moe Hardware Configs use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Moe Hardware Configs?

Skills that share tags, products or a category with Nemo Mbridge Perf Moe Hardware Configs: Add New Model (JakeATX/llamAmpere, 157 stars), Debug Cuda Crash (sgl-project/sglang, 37k stars), Code Review (JakeATX/llamAmpere, 157 stars) and Add Model (sohu-mptc/FlashRec, 107 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Moe Hardware Configs?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.