Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
MoE expert-parallel communication overlap in Megatron Bridge.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-comm-overlap --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-comm-overlap .claude/skills/nemo-mbridge-perf-moe-comm-overlap && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-moe-comm-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-comm-overlap into .claude/skills/nemo-mbridge-perf-moe-comm-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-comm-overlap", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-comm-overlapType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-comm-overlap --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-comm-overlap .agents/skills/nemo-mbridge-perf-moe-comm-overlap && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-moe-comm-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-comm-overlap into .agents/skills/nemo-mbridge-perf-moe-comm-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-comm-overlap", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-comm-overlap --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-comm-overlap .cursor/skills/nemo-mbridge-perf-moe-comm-overlap && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-moe-comm-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-comm-overlap into .cursor/skills/nemo-mbridge-perf-moe-comm-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-comm-overlap", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-moe-comm-overlap--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-comm-overlap --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-comm-overlap .gemini/skills/nemo-mbridge-perf-moe-comm-overlap && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-comm-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-comm-overlap into .gemini/skills/nemo-mbridge-perf-moe-comm-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-comm-overlap", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-comm-overlapInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-comm-overlap .github/skills/nemo-mbridge-perf-moe-comm-overlap && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-comm-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-comm-overlap into .github/skills/nemo-mbridge-perf-moe-comm-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-comm-overlap", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-comm-overlap --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-comm-overlap .opencode/skills/nemo-mbridge-perf-moe-comm-overlap && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-comm-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-comm-overlap into .opencode/skills/nemo-mbridge-perf-moe-comm-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-comm-overlap", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-moe-comm-overlapMoE expert-parallel communication overlap in Megatron Bridge.
Nemo Mbridge Perf Moe Comm Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. MoE expert-parallel communication overlap in Megatron Bridge. Covers dispatch/combine overlap, flex dispatcher backends, and expert wgrad scheduling.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Moe Comm Overlap loads about 1.9k tokens when it runs. Until then it costs about 46 tokens; SKILL.md has 771 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 771 words, ~1,873 tokens.
.claude/skills/nemo-mbridge-perf-moe-comm-overlap/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.For the higher-level overview, see:
Use MoE communication overlap when:
EP > 1Avoid turning it on as an early bring-up step. It is easier to validate after the dispatcher, routing mode, and recompute plan are already stable.
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
# Optional: delayed wgrad for additional overlap
cfg.comm_overlap.delay_wgrad_compute = True
# IMPORTANT: disable shared expert overlap when using dispatch overlap
cfg.model.moe_shared_expert_overlap = Falseexpert_model_parallel_size > 1num_moe_experts > 1moe_token_dispatcher_type must be "alltoall" or "flex"virtual_pipeline_model_parallel_size) must be set (non-None)Setting moe_flex_dispatcher_backend alone does not activate flex dispatch.
You must also set moe_token_dispatcher_type = "flex".
delay_wgrad_compute adds further constraints if CUDA-graph scopes include
attention or MoE-router work.A 2026-07-25 controlled Qwen3 30B-A3B pretraining comparison used 16 H100
GPUs, BF16, sequence length 4096, TP=1, PP=1, CP=1, EP=16,
MBS=1, GBS=1024, forced-balanced routing, HybridEP, and Transformer
Engine CUDA-graph scopes moe_router and moe_preprocess. The only
performance change was plain EP overlap; delayed wgrad stayed disabled.
| Case | Steady window | Step time | Model TFLOPS/GPU |
|---|---|---|---|
| EP overlap off | iterations 5-20 | 24.7138s | 244.039 |
| EP overlap on, search run | iterations 5-20 | 21.0725s | 286.208 |
| EP overlap on, independent validation | iterations 41-50 | 20.9920s | 287.305 |
The independent result reduced step time by 15.059% and increased throughput by 17.729% over the reproduced baseline. Loss remained finite, no iterations were skipped or NaN, and rank-0 peak allocated memory was 62.166 GiB.
A same-method rank-0 Nsight Systems comparison captured 463,348 kernels in each case:
| Profile metric | Overlap off | Overlap on |
|---|---|---|
| Communication concurrent with GEMM/attention | 9.079ms | 3,958.997ms |
| Communication time hidden by compute | 0.11% | 36.55% |
| GPU-active interval union | 22.821s | 21.221s |
| HybridEP dispatch-with-permute NVTX | 4.253s | 1.767s |
| HybridEP metadata-preprocess NVTX | 3.109s | 0.670s |
This is direct evidence that the gain came from hiding exposed HybridEP dispatch/combine work, not from changing the dispatcher, routing, graph scopes, batch shape, or parallel layout.
A 2026-05-18 current-main H100 x16 smoke on Qwen3 30B-A3B mock pretraining
used EP=16, alltoall, global batch size 1024, CUDA graphs disabled, and
moe_permute_fusion=false because the PyTorch 25.11 / TE / Triton stack failed
in Transformer Engine fused permutation in prior bring-up.
Results were directional rather than release-grade:
delay_wgrad_compute: 31.20s steady-state mean over
iterations 3-8Treat this as evidence that EP overlap can help an inter-node alltoall MoE
shape when communication is exposed. It is not proof that delayed wgrad is a
separate win, and it does not validate the fused permutation path. An earlier
2026-05-16 short smoke on the same shape showed the same pattern.
src/megatron/bridge/training/comm_overlap.pysrc/megatron/bridge/training/flex_dispatcher_backend.pysrc/megatron/bridge/training/config.pytests/unit_tests/training/test_comm_overlap.pytests/unit_tests/training/test_deepep.pyShared expert overlap conflict: moe_shared_expert_overlap and
overlap_moe_expert_parallel_comm can conflict. Disable shared expert
overlap when using the dispatch overlap path.
PP without VPP: MoE overlap requires VPP when pipeline parallelism is active. Without it, the overlap scheduling cannot interleave correctly.
Flex != backend flag: moe_flex_dispatcher_backend="deepep" alone
does nothing if moe_token_dispatcher_type is still "alltoall".
Conservative recipe defaults: Most public recipes leave MoE overlap disabled. You need to explicitly enable it via overrides.
Performance gains are workload-dependent: overlap helps most when dispatch communication is already a visible slice of step time. It is not guaranteed to help every small or lightly loaded EP run.
Summed kernel time is not wall time: concurrent kernels can run longer because they contend for SMs or bandwidth, so overlap may increase summed per-stream kernel duration while reducing the exposed interval union and end-to-end step time.
Look for overlap-related log messages during initialization. The comm overlap
validation in comm_overlap.py will raise if prerequisites are not met, so a
clean startup confirms the feature is active.
For a short performance-harness smoke, keep the command shape explicit and vary only one overlap knob at a time:
uv run python scripts/performance/run_script.py \
-m qwen \
-mr qwen3_30b_a3b \
--task pretrain \
-g h100 \
-c bf16 \
-ng 16 \
-gn 8 \
--max_steps 8 \
--cuda_graph_impl none \
--moe_flex_dispatcher_backend None \
--moe_a2a_overlap false \
--tokenizer_type NullTokenizer \
comm_overlap.overlap_moe_expert_parallel_comm=true \
comm_overlap.delay_wgrad_compute=false \
model.moe_shared_expert_overlap=falseIf fused MoE permutation fails during bring-up, add
model.moe_permute_fusion=false to separate overlap timing from runtime-stack
validation, then retest with the matched production container.
For performance validation, use an unprofiled steady window as the acceptance metric. Use a matched Nsight A/B to establish causality:
overlap_moe_expert_parallel_comm; keep
delay_wgrad_compute=false for the first isolation.Last signature refresh: 2026-08-03.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-comm-overlap of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Nemo Mbridge Perf Moe Comm Overlap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Moe Comm Overlap this skillNVIDIA/skills | 3.6k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.8k | Automated safety check: Pass | None | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
LMIXR/CV_Deployment_skill
基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
MoE expert-parallel communication overlap in Megatron Bridge. Nemo Mbridge Perf Moe Comm Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. MoE expert-parallel communication overlap in Megatron Bridge.
Nemo Mbridge Perf Moe Comm Overlap fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-comm-overlap in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-comm-overlap in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-comm-overlap in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-comm-overlap in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-comm-overlap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-comm-overlap, .gemini/skills/nemo-mbridge-perf-moe-comm-overlap, .github/skills/nemo-mbridge-perf-moe-comm-overlap and .opencode/skills/nemo-mbridge-perf-moe-comm-overlap in your project.
Going by SKILL.md and its folder, Nemo Mbridge Perf Moe Comm Overlap needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Moe Comm Overlap is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Moe Comm Overlap: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.