Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-expert-parallel-overlap --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-expert-parallel-overlap .claude/skills/nemo-mbridge-perf-expert-parallel-overlap && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-expert-parallel-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-expert-parallel-overlap into .claude/skills/nemo-mbridge-perf-expert-parallel-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-expert-parallel-overlap", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-expert-parallel-overlapType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-expert-parallel-overlap --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-expert-parallel-overlap .agents/skills/nemo-mbridge-perf-expert-parallel-overlap && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-expert-parallel-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-expert-parallel-overlap into .agents/skills/nemo-mbridge-perf-expert-parallel-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-expert-parallel-overlap", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-expert-parallel-overlap --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-expert-parallel-overlap .cursor/skills/nemo-mbridge-perf-expert-parallel-overlap && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-expert-parallel-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-expert-parallel-overlap into .cursor/skills/nemo-mbridge-perf-expert-parallel-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-expert-parallel-overlap", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-expert-parallel-overlap--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-expert-parallel-overlap --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-expert-parallel-overlap .gemini/skills/nemo-mbridge-perf-expert-parallel-overlap && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-expert-parallel-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-expert-parallel-overlap into .gemini/skills/nemo-mbridge-perf-expert-parallel-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-expert-parallel-overlap", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-expert-parallel-overlapInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-expert-parallel-overlap .github/skills/nemo-mbridge-perf-expert-parallel-overlap && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-expert-parallel-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-expert-parallel-overlap into .github/skills/nemo-mbridge-perf-expert-parallel-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-expert-parallel-overlap", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-expert-parallel-overlap --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-expert-parallel-overlap .opencode/skills/nemo-mbridge-perf-expert-parallel-overlap && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-expert-parallel-overlap" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-expert-parallel-overlap into .opencode/skills/nemo-mbridge-perf-expert-parallel-overlap/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-expert-parallel-overlap", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-expert-parallel-overlapValidate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.
Nemo Mbridge Perf Expert Parallel Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Expert Parallel Overlap loads about 3.5k tokens when it runs. Until then it costs about 61 tokens; SKILL.md has 1,186 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,186 words, ~3,536 tokens.
.claude/skills/nemo-mbridge-perf-expert-parallel-overlap/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Expert-parallel (EP) overlap hides the cost of token dispatch/combine all-to-all
communication by running it concurrently with expert FFN compute. Optionally,
delayed expert weight-gradient computation (delay_wgrad_compute) provides
additional overlap by deferring wgrad to overlap with the next layer's forward.
Bridge supports two dispatcher paths:
| Dispatcher | Backend | When to use |
|---|---|---|
alltoall | Standard MoE all-to-all | Default, broadest compatibility |
flex | DeepEP or HybridEP | Higher overlap on Ampere/Hopper/Blackwell |
Use EP overlap when:
EP > 1Prefer:
alltoall dispatcher for the first rollout (broader compatibility)flex + DeepEP/HybridEP when running on supported GPUs and seeking
additional gainsAvoid EP overlap when:
moe_shared_expert_overlap is enabledExpected outcome:
For the plain EP-overlap isolation benchmark, keep flex dispatch and delayed
wgrad disabled. The measured shape was Qwen3 MoE 30B-A3B SFT on 16 H100 GPUs:
EP=16, alltoall, BF16, global batch size 1024, CUDA graphs disabled,
moe_permute_fusion=false, measured over iterations 3-8.
Use these overrides for the plain-overlap case:
--cuda_graph_impl none \
--moe_flex_dispatcher_backend None \
--moe_a2a_overlap false \
comm_overlap.overlap_moe_expert_parallel_comm=true \
comm_overlap.delay_wgrad_compute=false \
model.moe_shared_expert_overlap=falseDo not use --moe_a2a_overlap true for this isolation test: the performance
harness helper enables both overlap_moe_expert_parallel_comm and
delay_wgrad_compute, so it does not isolate plain EP overlap.
Steady-window timing from that benchmark:
| Case | Steady mean | Relative |
|---|---|---|
| no EP overlap | 41.25s | 1.000x |
| EP overlap | 31.31s | 1.317x |
EP overlap plus delay_wgrad_compute | 31.20s | 1.322x |
This is evidence for enabling plain EP overlap on this inter-node all-to-all shape. It does not show a meaningful independent win from delayed wgrad, and it does not validate fused MoE permutation because that path was disabled for the runtime stack.
A 2026-07-25 controlled Qwen3 30B-A3B pretraining comparison validated plain EP overlap with the production HybridEP path:
Hardware: 16×H100
Precision: BF16
Sequence: 4096
Parallelism: TP1 / PP1 / CP1 / EP16
Batch: MBS1 / GBS1024
Routing: force balance
Dispatcher: flex + HybridEP
CUDA graph: Transformer Engine scopes moe_router + moe_preprocess
Delayed wgrad: disabled| Case | Steady window | Step time | Model TFLOPS/GPU |
|---|---|---|---|
| overlap off | iterations 5-20 | 24.7138s | 244.039 |
| overlap on, search run | iterations 5-20 | 21.0725s | 286.208 |
| overlap on, independent validation | iterations 41-50 | 20.9920s | 287.305 |
The independent run reduced step time by 15.059% and raised throughput by 17.729% over the reproduced baseline. Loss was finite, skipped and NaN iterations remained zero, and rank-0 peak allocated memory was 62.166 GiB.
A matched Nsight Systems comparison captured the same 463,348 rank-0 kernels per case. Enabling overlap increased communication concurrent with GEMM and attention from 9.079ms (0.11% of communication time) to 3,958.997ms (36.55%). GPU-active interval union fell from 22.821s to 21.221s.
Use this as evidence for the mechanism, not as a universal speedup promise. The dispatcher, graph scopes, routing, parallelism, batch shape, and runtime were held fixed while only plain EP overlap changed.
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.moe_shared_expert_overlap = False
cfg.model.expert_model_parallel_size = 8
cfg.model.num_moe_experts = 64
cfg.model.moe_token_dispatcher_type = "alltoall"
cfg.model.bf16 = True
cfg.model.fp16 = FalseEnable delay_wgrad_compute=True only after the plain overlap path is known to
work and its extra compatibility constraints have been checked.
from megatron.bridge.training.flex_dispatcher_backend import apply_flex_dispatcher_backend
cfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.moe_shared_expert_overlap = False
apply_flex_dispatcher_backend(cfg.model, moe_flex_dispatcher_backend="deepep")
# or: apply_flex_dispatcher_backend(cfg.model, moe_flex_dispatcher_backend="hybridep")Benchmark plain EP overlap first. Enable delay_wgrad_compute=True only as a
separate follow-up A/B after its CUDA-graph and TE compatibility constraints
are satisfied.
expert_model_parallel_size > 1num_moe_experts > 1moe_token_dispatcher_type must be "alltoall" or "flex"moe_shared_expert_overlap = False>= 2.6.0PP > 1, virtual_pipeline_model_parallel_size must be setrecompute_granularity != "full", recompute_method = None,
recompute_num_layers = Nonemtp_num_layers must be None or 1delay_wgrad_compute requires overlap_moe_expert_parallel_comm as a
prerequisitedelay_wgrad_compute with overlap_grad_reduce requires TE >= 2.7.0delay_wgrad_compute with gradient_accumulation_fusion requires TE >= 2.7.0attn scope + delay_wgrad_compute requires TE >= 2.12.0,
gradient_accumulation_fusion = True, and no attention biascfg.comm_overlap.overlap_moe_expert_parallel_comm = True
cfg.comm_overlap.delay_wgrad_compute = False
cfg.model.expert_model_parallel_size = 4
cfg.model.num_moe_experts = 64
cfg.model.moe_token_dispatcher_type = "alltoall"
cfg.model.moe_shared_expert_overlap = False
cfg.model.bf16 = TrueUse this as the correctness-first starting point. Add delayed wgrad, flex dispatch, and CUDA-graph interactions only after the plain overlap path is known to work.
Performance harness example inside a Slurm allocation. Keep the model, parallelism, dispatcher, and runtime fixed, and vary only the two overlap overrides:
uv run python scripts/performance/run_script.py \
-m qwen \
-mr qwen3_30b_a3b \
--task pretrain \
-g h100 \
-c bf16 \
-ng 16 \
-gn 8 \
--max_steps 8 \
--cuda_graph_impl none \
--moe_flex_dispatcher_backend None \
--moe_a2a_overlap false \
--tokenizer_type NullTokenizer \
comm_overlap.overlap_moe_expert_parallel_comm=true \
comm_overlap.delay_wgrad_compute=false \
model.moe_shared_expert_overlap=falseDo not use --moe_a2a_overlap true when separating plain EP overlap from
delayed wgrad: the performance harness helper enables both
overlap_moe_expert_parallel_comm and delay_wgrad_compute.
Unit test verification:
uv run python -m pytest \
tests/unit_tests/training/test_comm_overlap.py -k "moe" \
tests/unit_tests/training/test_deepep.py -quv run python -m pytest \
tests/unit_tests/training/test_comm_overlap.py \
tests/unit_tests/training/test_deepep.py -qAfter a successful run with EP overlap:
CommOverlapConfig finalizationoverlap_moe_expert_parallel_comm appears as True in the logged
configmoe_token_dispatcher_type = "flex" and
the correct backend in logsUse an unprofiled steady window for the throughput acceptance result. Use a matched profile to explain the mechanism:
if self.user_comm_overlap_cfg.overlap_moe_expert_parallel_comm is True:
assert model_cfg.expert_model_parallel_size > 1, ...
assert model_cfg.num_moe_experts > 1, ...
assert model_cfg.moe_token_dispatcher_type in ["alltoall", "flex"], ...
assert model_cfg.bf16 or model_cfg.fp16, ...
assert is_torch_min_version("2.6.0"), ...
# ... PP + VPP check, recompute checks, shared_expert_overlap check ...if self.user_comm_overlap_cfg.delay_wgrad_compute is True:
# TE version checks for overlap_grad_reduce and gradient_accumulation_fusion
# CUDA graph scope validations for delayed wgrad
assert overlap_moe_expert_parallel_comm, ...def apply_flex_dispatcher_backend(...):
# GPU architecture check for DeepEP / HybridEP
model_config.moe_token_dispatcher_type = "flex"
model_config.moe_flex_dispatcher_backend = moe_flex_dispatcher_backend
model_config.moe_shared_expert_overlap = Falsedef _set_moe_a2a_overlap_overrides(recipe, moe_a2a_overlap=False):
if moe_a2a_overlap:
recipe.comm_overlap.overlap_moe_expert_parallel_comm = True
recipe.comm_overlap.delay_wgrad_compute = True
recipe.model.moe_shared_expert_overlap = False| File | Coverage |
|---|---|
tests/unit_tests/training/test_comm_overlap.py | EP overlap validation, delayed wgrad, CUDA graph + wgrad interaction |
tests/unit_tests/training/test_deepep.py | DeepEP/HybridEP helper activation and GPU gating |
| Symptom | Likely Cause | How To Confirm | Fix |
|---|---|---|---|
assert expert_model_parallel_size > 1 | EP not configured | Check expert_model_parallel_size | Set EP > 1 |
assert moe_token_dispatcher_type | Wrong dispatcher | Check dispatcher type | Use "alltoall" or "flex" |
| assert on BF16/FP16 | Wrong precision | Check bf16 and fp16 | Set bf16 = True |
| hang during training | PyTorch < 2.6 | Check PyTorch version | Upgrade to >= 2.6.0 |
assert virtual_pipeline_model_parallel_size | PP > 1 without VPP | Check PP and VPP config | Set VPP when PP > 1 |
assert recompute_granularity | Full recompute enabled | Check recompute settings | Disable full recompute |
assert overlap_moe_expert_parallel_comm required | delayed wgrad without EP overlap | Check delay_wgrad_compute without overlap | Enable EP overlap first |
assert gradient_accumulation_fusion | CUDA graph + delayed wgrad | Check graph scope + wgrad settings | Enable gradient_accumulation_fusion |
| assert on attention bias | CUDA graph attn + delayed wgrad + bias | Check add_bias_linear / add_qkv_bias | Disable attention bias |
| no throughput gain from flex dispatcher | apply_flex_dispatcher_backend not called | Check moe_token_dispatcher_type in logs | Call apply_flex_dispatcher_backend(...) |
| DeepEP/HybridEP silently skipped | Unsupported GPU | Check warning logs | Run on Ampere/Hopper/Blackwell |
| summed kernel time increases after overlap | Expected concurrency contention or a regression | Compare interval unions, comm/compute intersection, and unprofiled step time | Judge overlap from exposed wall time, not summed per-stream duration |
moe_flex_dispatcher_backend alone does not activate flex dispatch —
you must call apply_flex_dispatcher_backend(...).Last signature refresh: 2026-08-03.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-expert-parallel-overlap of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
Nemo Mbridge Perf Expert Parallel Overlap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Expert Parallel Overlap this skillNVIDIA/skills | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 925 | — | ~2.8k | Automated safety check: Pass | None | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
LMIXR/CV_Deployment_skill
基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP. Nemo Mbridge Perf Expert Parallel Overlap is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.
Nemo Mbridge Perf Expert Parallel Overlap fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-expert-parallel-overlap in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-expert-parallel-overlap in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-expert-parallel-overlap in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-expert-parallel-overlap in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-expert-parallel-overlap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-expert-parallel-overlap, .gemini/skills/nemo-mbridge-perf-expert-parallel-overlap, .github/skills/nemo-mbridge-perf-expert-parallel-overlap and .opencode/skills/nemo-mbridge-perf-expert-parallel-overlap in your project.
Going by SKILL.md and its folder, Nemo Mbridge Perf Expert Parallel Overlap needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Expert Parallel Overlap is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Expert Parallel Overlap: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.