Official agent skill

Nemo Mbridge Perf Tp Dp Comm Overlap

by NVIDIA in NVIDIA/skills

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Perf Tp Dp Comm Overlap

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-tp-dp-comm-overlap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-tp-dp-comm-overlap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-tp-dp-comm-overlap .claude/skills/nemo-mbridge-perf-tp-dp-comm-overlap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-tp-dp-comm-overlap
GitHub stars
3.6k
Token cost
~924 tokens
SKILL.md length
145 words
Files
6
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

  • Works in 5 steps: TP overlap silently disables itself if… → PP overlap is not enabled for all PP… → bucket_size is a parameter-count knob,… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Enablement, Code Anchors, Pitfalls and Verification
  • Calls uv

What it does

Nemo Mbridge Perf Tp Dp Comm Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/nemo-mbridge-perf-tp-dp-comm-overlap”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. TP overlap silently disables itself if sequence_parallel=False or Transformer Engine is unavailable.
  2. PP overlap is not enabled for all PP cases. Bridge only auto-selects overlap_p2p_comm=True when PP > 1 and VPP > 1.
  3. bucket_size is a parameter-count knob, not a byte-size knob.
  4. grad_reduce_in_fp32 and fp8_param_gather should be set through mixed precision, not as standalone DDP tuning first.
  5. CUDA_DEVICE_MAX_CONNECTIONS and LayerNorm SM margin are launch-time plugin settings, not CommOverlapConfig fields.

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Tp Dp Comm Overlap loads about 924 tokens when it runs. Until then it costs about 48 tokens; SKILL.md has 145 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~48
When it runs · the whole SKILL.md, loaded when a task matches
~924

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 145 words, ~924 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-tp-dp-comm-overlap/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-tp-dp-comm-overlap
description
Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
license
Apache-2.0
when_to_use
Enabling TP/DP/PP comm overlap, or tracing a throughput regression to a comm overlap config change; 'overlap_param_gather', 'overlap_grad_reduce'…

TP / DP / PP Communication Overlap Skill

For stable background and recommendation level, see:

  • @docs/training/communication-overlap.md

Enablement

Minimal Bridge override:

python
from megatron.bridge.training.comm_overlap import CommOverlapConfig

cfg.model.tensor_model_parallel_size = 4
cfg.model.sequence_parallel = True
cfg.model.pipeline_model_parallel_size = 4
cfg.model.virtual_pipeline_model_parallel_size = 2

cfg.comm_overlap = CommOverlapConfig(
    tp_comm_overlap=True,
)

cfg.ddp.use_distributed_optimizer = True
cfg.ddp.overlap_grad_reduce = True
cfg.ddp.overlap_param_gather = True

Optional TP preset:

python
from megatron.bridge.training.comm_overlap import userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048

cfg.comm_overlap.tp_comm_overlap_cfg = userbuffers_bf16_h100_h12288_tp4_mbs1_seqlen2048

Precision knobs belong to mixed precision:

python
cfg.mixed_precision.grad_reduce_in_fp32 = False
cfg.mixed_precision.fp8_param_gather = False

Code Anchors

Bridge overlap gating:

439449srcmegatronbri
if self.user_comm_overlap_cfg.tp_comm_overlap is True:
    if model_cfg.tensor_model_parallel_size < 2:
        ...
    elif not model_cfg.sequence_parallel:
        ...
    elif not HAVE_TE:
        ...

PP overlap selection:

451458srcmegatronbri
if model_cfg.pipeline_model_parallel_size > 1:
    if vp_size > 1:
        comm_overlap_cfg.overlap_p2p_comm = True
        comm_overlap_cfg.batch_p2p_comm = False
    else:
        comm_overlap_cfg.overlap_p2p_comm = False
        comm_overlap_cfg.batch_p2p_comm = True

DP overlap defaults:

572579srcmegatronbri
if self.data_parallel_size > 1:
    comm_overlap_cfg.bucket_size = 128 * 1024 * 1024
    comm_overlap_cfg.overlap_grad_reduce = True
    comm_overlap_cfg.overlap_param_gather = True

Launch-time env tuning:

570609srcmegatronbri
executor.env_vars["CUDA_DEVICE_MAX_CONNECTIONS"] = str(cuda_device_max_connections)
...
executor.env_vars["NVTE_FWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)
executor.env_vars["NVTE_BWD_LAYERNORM_SM_MARGIN"] = str(self.layernorm_sm_margin)

Pitfalls

  1. TP overlap silently disables itself if sequence_parallel=False or Transformer Engine is unavailable.
  2. PP overlap is not enabled for all PP cases. Bridge only auto-selects overlap_p2p_comm=True when PP > 1 and VPP > 1.
  3. bucket_size is a parameter-count knob, not a byte-size knob.
  4. grad_reduce_in_fp32 and fp8_param_gather should be set through mixed precision, not as standalone DDP tuning first.
  5. CUDA_DEVICE_MAX_CONNECTIONS and LayerNorm SM margin are launch-time plugin settings, not CommOverlapConfig fields.

Verification

Use the checked-in overlap unit coverage first:

bash
uv run python -m pytest tests/unit_tests/training/test_comm_overlap.py -q

Optional second check if nemo_run is available:

bash
uv run python -m pytest tests/unit_tests/recipes/test_run_plugins.py -q

Success criteria:

  • first command reports 26 passed
  • second command validates plugin-owned env wiring when not skipped

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-tp-dp-comm-overlap of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Nemo Mbridge Perf Tp Dp Comm Overlap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Tp Dp Comm Overlap compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Tp Dp Comm Overlap this skillNVIDIA/skills3.6k—~924Automated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    188 GitHub stars~547 tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemo Mbridge Perf Tp Dp Comm Overlap

What does Nemo Mbridge Perf Tp Dp Comm Overlap do?

Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. Nemo Mbridge Perf Tp Dp Comm Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Operational guide for enabling TP, DP, and PP communication overlap in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

When should I use Nemo Mbridge Perf Tp Dp Comm Overlap?

Nemo Mbridge Perf Tp Dp Comm Overlap fits situations like: AI & LLM Engineering work in your project.

How do I install Nemo Mbridge Perf Tp Dp Comm Overlap in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-tp-dp-comm-overlap -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-tp-dp-comm-overlap in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-tp-dp-comm-overlap in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Tp Dp Comm Overlap in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-tp-dp-comm-overlap -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-tp-dp-comm-overlap in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-tp-dp-comm-overlap in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Tp Dp Comm Overlap in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-tp-dp-comm-overlap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-tp-dp-comm-overlap, .gemini/skills/nemo-mbridge-perf-tp-dp-comm-overlap, .github/skills/nemo-mbridge-perf-tp-dp-comm-overlap and .opencode/skills/nemo-mbridge-perf-tp-dp-comm-overlap in your project.

What does Nemo Mbridge Perf Tp Dp Comm Overlap need to run?

Going by SKILL.md and its folder, Nemo Mbridge Perf Tp Dp Comm Overlap needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Perf Tp Dp Comm Overlap access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Tp Dp Comm Overlap safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Tp Dp Comm Overlap use?

Nemo Mbridge Perf Tp Dp Comm Overlap is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Tp Dp Comm Overlap use?

About 924 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Tp Dp Comm Overlap?

Skills that share tags, products or a category with Nemo Mbridge Perf Tp Dp Comm Overlap: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Tp Dp Comm Overlap?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.