Official agent skill

Nemo Mbridge Perf Moe Long Context

by NVIDIA in NVIDIA/skills

Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Perf Moe Long Context

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-long-context -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-long-context --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-long-context .claude/skills/nemo-mbridge-perf-moe-long-context && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-moe-long-context
GitHub stars
3.5k
Token cost
~1.2k tokens
SKILL.md length
569 words
Files
6
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills.

  • Works in 6 steps: Start from a 4K shard target: a good… → Keep DP alive if possible: long-context… → Prefer selective recompute: recompute… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers What Changes At Long Context, Rounded Scaling Patterns, CP Sizing Rules Of Thumb and Representative Config Families, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Nemo Mbridge Perf Moe Long Context is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Long-context MoE training guidance for Megatron Bridge. Covers CP sizing, selective recompute, dispatcher choices, and practical patterns from DSV3, Qwen3, and Qwen3-Next long-context experiments.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering. It works with Qwen, NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/nemo-mbridge-perf-moe-long-context”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Start from a 4K shard target: a good first guess is
  2. Keep DP alive if possible: long-context scaling becomes brittle once CP,
  3. Prefer selective recompute: recompute modules such as up_proj, norm,
  4. Avoid SDPA-heavy recompute at very long context: recomputing attention
  5. Use TP as another lever on NVL72 systems: GB200 and GB300 runs can
  6. Assume GBS will need to shrink: as CP rises and DP falls, you may need

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Moe Long Context loads about 1.2k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 569 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 569 words, ~1,234 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-moe-long-context/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-moe-long-context
description
Long-context MoE training guidance for Megatron Bridge. Covers CP sizing, selective recompute, dispatcher choices, and practical patterns from DSV3, Qwen3, and Qwen3-Next long-context experiments.
license
Apache-2.0
when_to_use
Training MoE at long sequence lengths, or investigating a commit that caused long-context MoE OOM or degraded throughput; 'long context MoE', '128k tokens'…

MoE Long-Context Training

Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-long-context/card.yaml

What Changes At Long Context

Once sequence length moves well past the 4K-class regime, attention memory and activation residency become the dominant constraints. For MoE models, that usually means you need some combination of:

  • context parallelism
  • selective recompute
  • lower precision
  • CPU offload for optimizer state
  • a dispatcher and PP layout that do not waste the smaller remaining DP budget

Rounded Scaling Patterns

DSV3 on H100

The DSV3 long-context runs show a stable pattern:

  • selective recompute works better than full recompute once you move past the shortest contexts
  • throughput stays in a fairly narrow band from mid-length through very long contexts if CP is increased appropriately
  • the trade shifts from "memory fit" to "GPU-count feasibility" as CP grows

In other words, long context does not immediately collapse utilization if the layout is chosen well, but it does consume the DP budget very quickly.

Qwen3-Next on GB200

Qwen3-Next behaves more like a memory-sensitive medium-scale model:

  • 8K and 32K remain practical with moderate CP
  • 64K is possible, but the throughput drop is noticeable and memory becomes much tighter
  • pipeline layout and grouped-GEMM improvements matter almost as much as CP
Qwen3 235B on GB200

Qwen3 235B shows that long context can still be efficient on NVL72 systems when TP, CP, and HybridEP are coordinated. The best 128K-class configurations are not just "fit-only" recipes; they can remain highly efficient if routing, parallelism, and recompute are balanced.

CP Sizing Rules Of Thumb

  1. Start from a 4K shard target: a good first guess is CP ~= seq_len / 4096, then round to a practical power-of-two layout.

  2. Keep DP alive if possible: long-context scaling becomes brittle once CP, EP, TP, and PP together squeeze DP down to the floor.

  3. Prefer selective recompute: recompute modules such as up_proj, norm, moe, moe_act, or mlp before reaching for full recompute.

  4. Avoid SDPA-heavy recompute at very long context: recomputing attention internals can add a lot of work for less memory benefit than recomputing smaller MoE and MLP-side modules.

  5. Use TP as another lever on NVL72 systems: GB200 and GB300 runs can sometimes trade some CP for TP while still staying efficient.

  6. Assume GBS will need to shrink: as CP rises and DP falls, you may need to reduce global batch size or accept higher GA.

Show full SKILL.md (184 more words)Show less

Representative Config Families

DSV3 at 128K on H100
text
TP=1  CP=32  EP=32  PP=8  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
DSV3 at 256K on H100
text
TP=1  CP=64  EP=32  PP=8  EDP=2  VPP=4
Precision: FP8-class
Dispatcher: DeepEP
Recompute: up_proj, norm, moe, mlp
Extra memory help: optimizer CPU offload
Qwen3 235B at 128K on GB200
text
TP=4  CP=4  EP=32  PP=4  VPP=12
Precision: BF16 or MXFP8
Dispatcher: HybridEP
Recompute: moe_act, norm
CUDA Graph: attn + moe_router + moe_preprocess

Recompute And CUDA Graph Guidance

For long-context MoE training:

  • start with selective recompute
  • add CUDA graphs only after the shapes and routing path are stable
  • keep sequence length and MBS fixed when using CUDA graphs
  • if the run depends on highly dynamic batches, prefer eager execution

Useful references:

  • @docs/training/activation-recomputation.md
  • @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md

Pitfalls

  1. CP does not replace EP or PP: it adds another dimension; it does not make the others disappear.

  2. A good 4K baseline can still be a bad long-context baseline: routing mode, recompute choice, and offload strategy often need to change.

  3. GPU-count feasibility becomes the real constraint: very long context can look fine in a single recipe, then become impossible once EP and PP are added honestly across the full model.

  4. CUDA graphs need static shapes: variable-length batches and opportunistic padding strategies can silently break the path.

  5. Container and kernel support matters more at 128K+: long-context paths tend to rely on newer kernels and bug fixes than short-context bring-up does.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-long-context of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemo Mbridge Perf Moe Long Context next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Moe Long Context compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Moe Long Context this skillNVIDIA/skills3.5k—~1.2kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    170 GitHub stars~547 tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemo Mbridge Perf Moe Long Context

What does Nemo Mbridge Perf Moe Long Context do?

Long-context MoE training guidance for Megatron Bridge. An agent skill from NVIDIA/skills. Nemo Mbridge Perf Moe Long Context is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Long-context MoE training guidance for Megatron Bridge.

When should I use Nemo Mbridge Perf Moe Long Context?

Nemo Mbridge Perf Moe Long Context fits situations like: AI & LLM Engineering work in your project.

How do I install Nemo Mbridge Perf Moe Long Context in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-long-context -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-long-context in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-long-context in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Moe Long Context in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-long-context -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-long-context in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-long-context in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Moe Long Context in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-long-context -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-long-context, .gemini/skills/nemo-mbridge-perf-moe-long-context, .github/skills/nemo-mbridge-perf-moe-long-context and .opencode/skills/nemo-mbridge-perf-moe-long-context in your project.

What does Nemo Mbridge Perf Moe Long Context need to run?

SKILL.md names no scripts, command-line tools or credentials: Nemo Mbridge Perf Moe Long Context is instructions for the agent only.

Does Nemo Mbridge Perf Moe Long Context access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Moe Long Context safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Moe Long Context use?

Nemo Mbridge Perf Moe Long Context is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Moe Long Context use?

About 1.2k tokens (SKILL.md is roughly 4.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Moe Long Context?

Skills that share tags, products or a category with Nemo Mbridge Perf Moe Long Context: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Moe Long Context?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.