Official agent skill

Nemo Mbridge Perf Moe Optimization Workflow

by NVIDIA in NVIDIA/skills

Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

OfficialApache-2.0Auto-check passedDevelopment

Install Nemo Mbridge Perf Moe Optimization Workflow

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-optimization-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-optimization-workflow .claude/skills/nemo-mbridge-perf-moe-optimization-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-moe-optimization-workflow
GitHub stars
3.5k
Token cost
~3.3k tokens
SKILL.md length
1,656 words
Files
6
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

  • Works in 6 steps: Freeze The Measurement Contract → Make The Run Memory-Feasible → Choose Parallelism For Scale → …
  • Tasks that involve Performance optimization
  • SKILL.md covers Quick Reference, First Answer Checklist, Phase 0: Freeze The… and Phase 1: Make The Run…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Nemo Mbridge Perf Moe Optimization Workflow is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Evidence-gated workflow for MoE performance optimization in Megatron Bridge. Covers measurement contracts, the Three Walls framework, parallel folding, profiling, matched A/B tuning, and final validation.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in Development, covering Performance optimization. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Performance optimization

Example prompts

  • “/nemo-mbridge-perf-moe-optimization-workflow”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Freeze The Measurement Contract
  2. Make The Run Memory-Feasible
  3. Choose Parallelism For Scale
  4. Profile The Dominant Bottleneck
  5. Retune From Evidence
  6. Validate And Package Evidence

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Moe Optimization Workflow loads about 3.3k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 1,656 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,656 words, ~3,259 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-moe-optimization-workflow/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-moe-optimization-workflow
description
Evidence-gated workflow for MoE performance optimization in Megatron Bridge. Covers measurement contracts, the Three Walls framework, parallel folding, profiling, matched A/B tuning, and final validation.
license
Apache-2.0
when_to_use
Full MoE throughput tuning sweep, or diagnosing a MoE throughput regression after a commit or config change; 'optimize MoE throughput', 'MoE perf tuning'…

MoE Training Optimization Workflow

Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-optimization-workflow/card.yaml Source: Scalable Training of MoE Models with Megatron Core

Quick Reference

Start with the paper's Three Walls:

  • memory wall
  • communication wall
  • compute-efficiency wall

For operational diagnosis, split the compute-efficiency wall into compute and host/launch bottlenecks. They need different evidence and different fixes. MoE tuning is iterative, so use this order:

text
freeze the measurement contract -> fit -> scale -> profile -> retune -> validate

First Answer Checklist

For MoE optimization workflow prompts, present the response in this order:

  1. Freeze the measurement contract: record the exact model and task, hardware and topology, container and commits, data and routing semantics, precision, sequence and batch shape, parallelism, graph scopes, and the steady-state metric window. Label each candidate as training-equivalent or benchmark-only.
  2. Fit: make the model memory-feasible first. Use the smallest model parallelism that fits, prefer selective recompute before full recompute, add offloading only after recompute and parallelism are insufficient, and use --fake-init-process-group to sanity-check large layouts.
  3. Scale: maximize DP after the model fits, keep hot communication inside the fastest interconnect, use PP plus VPP for multi-node scaling, prefer EP over extra TP for expert layers, and add CP when long context makes attention memory dominant.
  4. Profile: identify the dominant wall: memory, communication, host overhead, or compute.
  5. Retune: change one variable at a time based on the profiled bottleneck. Dispatcher, overlap, lower precision, CUDA graphs, and recompute are candidates, not hardware defaults.
  6. Validate: use short matched screens to reject candidates, then run the winner for at least 50 steps. Verify the requested backend or graph replay actually ran, time a declared post-warmup window, and report loss health, skipped/NaN iterations, memory, step time, and model TFLOPS/GPU.
  7. Include the exact Parallel Folding meshes: Attention: TP x CP x DP x PP and MoE: ETP x EP x EDP x PP.
  8. Use alltoall for safe bring-up, then A/B flex + deepep and flex + hybridep when their packages and target topology support them. Start from BF16 and eager execution; introduce lower precision or the narrowest useful CUDA-graph scope only after profiling justifies it.

Phase 0: Freeze The Measurement Contract

A comparison is valid only when the following stay fixed unless they are the single variable under test:

  • Bridge, MCore, Transformer Engine, container, CUDA, and NCCL versions
  • GPU count, SKU, node topology, and launcher/environment settings
  • model, task, data path, sequence length, MBS, GBS, and optimizer settings
  • TP, PP, VPP, CP, EP, ETP, and DP layout
  • routing semantics, precision, recompute, dispatcher, overlap, and graph scope
  • warmup and steady-state timing windows

Separate two acceptance classes:

  • Training-equivalent changes preserve the intended routing, loss, data, optimizer, and checkpoint/resume behavior.
  • Benchmark-only changes such as forced load balancing are useful for controlled kernel studies, but cannot establish production-training correctness or convergence.

Phase 1: Make The Run Memory-Feasible

Start with a configuration that fits reliably before chasing throughput.

Recommended order:

  1. Use the smallest amount of model parallelism that still fits.
  2. Turn on selective recompute before falling back to full recompute.
  3. Add offloading only when recompute and parallelism are still insufficient.
  4. Use --fake-init-process-group to sanity-check large parallel layouts on a single GPU before burning cluster time.
Recompute guidance

Prefer selective recompute for MoE runs:

  • good first choices: layernorm, core_attn, moe_act, mlp, or model-specific modules (shared_experts, mla_up_proj)
  • use full recompute only when the run still does not fit
  • revisit recompute after enabling CUDA graphs, because some graph scopes and full recompute paths do not mix well

As a rule of thumb, fine-grained recompute often recovers most of the needed memory while keeping throughput much closer to the non-recompute baseline than full-layer recompute does.

Phase 2: Choose Parallelism For Scale

Priority order:

  1. Maximize DP once the model fits.
  2. Keep the hot communication path inside the fast interconnect when possible.
  3. Use PP, plus VPP if needed, for multi-node scaling.
  4. Prefer EP over extra TP for expert layers.
  5. Add CP for long context once sequence length makes attention memory dominant.
Parallel Folding

Parallel Folding decouples attention and MoE parallelism so you do not have to pick a single compromise layout:

text
Attention: TP × CP × DP × PP
MoE:       ETP × EP × EDP × PP

Key knobs:

  • --expert-model-parallel-size
  • --expert-tensor-parallel-size

Use it when attention prefers some TP or CP, but expert layers benefit from a larger EP degree than the dense layers can tolerate.

Phase 3: Profile The Dominant Bottleneck

BottleneckWhat it looks likePrimary fixes
MemoryRun fits only with aggressive full recompute or OOMs during warmupselective recompute, FP8, offloading, better PP layout
CommunicationNsight shows large all-to-all or collective blocksDeepEP or HybridEP, EP overlap, DP/TP overlap, better PP layout
Host overheadGPU gaps, launch-bound traces, Python overheadCUDA graphs, --manual-gc, higher MBS, CPU affinity tuning
ComputeLow SM utilization after comm and host issues are addressedgrouped GEMM, fusion work, FP8, dispatcher-specific kernel tuning
Profile overlap without misreading it

Use unprofiled steady iterations for the acceptance metric and a matched profile for causal explanation:

  1. Change one overlap or dispatcher variable at a time; keep routing, graph scopes, parallelism, batch shape, and runtime fixed.
  2. Build interval unions for communication kernels and compute kernels, then measure their intersection to quantify hidden communication.
  3. Do not add kernel durations and call the result wall time. Concurrent kernels may run longer because of SM or bandwidth contention even while the exposed GPU-active union and end-to-end step time fall.
  4. Corroborate the trace with dispatch/combine NVTX ranges, steady step time, model TFLOPS/GPU, loss finiteness, skipped/NaN counts, and peak memory.

On a controlled 16×H100 Qwen3 30B-A3B HybridEP run, plain EP overlap increased communication hidden by GEMM/attention from 0.11% to 36.55%. The unprofiled step fell from 24.7138s to 20.9920s and throughput rose from 244.039 to 287.305 model TFLOPS/GPU. delay_wgrad_compute remained disabled.

Phase 4: Retune From Evidence

Choose the smallest candidate that targets the profiled bottleneck and change one variable at a time.

Show full SKILL.md (700 more words)Show less
Dispatcher And Overlap Guidance

Use dispatcher choice as a bottleneck fix, not as a hardware lookup table.

  • moe_token_dispatcher_type="alltoall": safest bring-up path, fine for smaller EP sizes
  • moe_token_dispatcher_type="flex" + moe_flex_dispatcher_backend="deepep": candidate when DeepEP is installed and communication is exposed
  • moe_token_dispatcher_type="flex" + moe_flex_dispatcher_backend="hybridep": topology-sensitive candidate on both NVL8 and NVL72 systems when HybridEP is installed

HybridEP plus plain EP overlap is the current measured winner for the canonical 16×H100 Qwen3 30B-A3B shape, while the canonical 256×H100 Qwen3 235B recipe uses standard alltoall plus overlap. Benchmark backend compatibility and throughput in the target container; neither GPU name nor EP degree determines the winner by itself.

If the all-to-all path is visible in profiles, combine dispatcher tuning with:

  • --overlap-moe-expert-parallel-comm
  • --overlap-grad-reduce
  • --tp-comm-overlap

Test plain EP overlap, shared-expert overlap, and delayed weight-gradient compute as separate candidates first. A combination can regress even when one component helped on another model.

Lower-Precision Candidate Matrix

Start with a verified BF16 baseline. Hardware capability only determines which lower-precision candidates are legal; it does not guarantee a speedup.

PlatformCandidate after BF16 is stable
Hopperper-tensor, current-scaling, or blockwise FP8 supported by the target stack
BlackwellMXFP8 or another supported FP8 recipe
Blackwell, speed-first explorationNVFP4 after the BF16/FP8 path is stable

Keep the router in FP32. The largest wins usually come from expert GEMMs and other heavy matrix math, not from trying to quantize every small MoE component. Require logs or traces showing that the intended kernels ran, and judge the candidate by end-to-end steady step time rather than theoretical peak FLOPS.

CUDA Graphs For MoE

Use CUDA graphs only after a profile shows meaningful host/launch gaps. For dropless MoE, start with the narrowest partial TE-scoped graph candidate:

  • moe_router
  • moe_preprocess

Add attn only if it is supported for the model and improves the same matched stack. A successful capture is not evidence of a speedup, and a graph win can disappear after dispatcher, overlap, or precision changes.

This path keeps dynamic expert work outside the graph. Budget extra memory, verify that shapes remain static, confirm replay rather than capture alone, and time only post-capture iterations.

Use full-iteration graphs only for graph-friendly workloads such as drop-and-pad or tightly controlled static-shape experiments.

Related references:

  • @skills/nemo-mbridge-perf-cuda-graphs/SKILL.md
  • @docs/training/cuda-graphs.md
  • @docs/training/activation-recomputation.md

Phase 5: Validate And Package Evidence

Use 6–12 post-warmup iterations for inexpensive screening when the workload allows it. For the selected candidate, run at least 50 steps and report a fixed steady window such as steps 41–50. The final evidence bundle should contain:

  • exact command/config diff, commits, container, hardware, and topology
  • declared routing/data semantics and training-equivalent vs benchmark-only label
  • proof that the intended dispatcher, precision kernels, overlap, and graph replay were active
  • step time and model TFLOPS/GPU from the same unprofiled steady window
  • finite loss, skipped/NaN counts, peak memory, and checkpoint/optimizer-state validation when the production path requires it
  • matched A/B profile evidence for the claimed causal mechanism

Do not attribute the total gain of a final multi-change winner to one earlier A/B. For example, the Qwen3 overlap experiment isolated a rise from 244.039 to 287.305 TFLOPS/GPU; the later canonical recipe reached 299.352 after additional HybridEP tuning. They answer different questions.

Pitfalls

  1. Do not optimize in the wrong order: fitting the model and selecting sane parallelism matter more than micro-optimizations.

  2. Platform changes the limiting wall: H100-class runs often feel more communication-bound, while GB200 or GB300 runs often expose CPU or launch overhead earlier.

  3. FP8 MFU can look misleadingly low: compare absolute throughput as well as MFU when switching precision modes.

  4. CUDA graphs and recompute interact: TE-scoped graphs are usually paired with selective recompute, not blanket full recompute.

  5. Parallel Folding is not optional at large scale: once attention and expert layers want clearly different layouts, a single shared TP or EP plan becomes a tax on both.

  6. Summed kernel time is not exposed time: use interval unions and communication/compute intersection when validating overlap.

  7. Benchmark-only semantics are not production acceptance: forced routing, synthetic data, or disabled optimizer/checkpoint paths must be disclosed and validated separately from training-equivalent results.

  8. Feature activation needs evidence: a config dump is insufficient when a backend can fall back, a graph can capture without helping, or a lower- precision recipe can miss the intended kernels.

Last signature refresh: 2026-08-03.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-optimization-workflow of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Nemo Mbridge Perf Moe Optimization Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Moe Optimization Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Moe Optimization Workflow this skillNVIDIA/skills3.5k—~3.3kAutomated safety check: PassApache-2.0
Tilelang SkillslowlyC/agent-gpu-skills169—~1.8kAutomated safety check: PassMIT
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone
Cuda Profilingmohitmishra786/low-level-dev-skills253—~1.6kAutomated safety check: NotesMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0

Similar skills

  • Tilelang Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

    169 GitHub stars~1.8k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Cuda Profiling

    mohitmishra786/low-level-dev-skills

    CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 2 days ago
    DevelopmentAuto-check: notes

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Categories

Questions about Nemo Mbridge Perf Moe Optimization Workflow

What does Nemo Mbridge Perf Moe Optimization Workflow do?

Evidence-gated workflow for MoE performance optimization in Megatron Bridge. Nemo Mbridge Perf Moe Optimization Workflow is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Evidence-gated workflow for MoE performance optimization in Megatron Bridge.

When should I use Nemo Mbridge Perf Moe Optimization Workflow?

Nemo Mbridge Perf Moe Optimization Workflow fits situations like: tasks that involve Performance optimization.

How do I install Nemo Mbridge Perf Moe Optimization Workflow in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-optimization-workflow in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-optimization-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Moe Optimization Workflow in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-optimization-workflow in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-optimization-workflow in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Moe Optimization Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-optimization-workflow, .gemini/skills/nemo-mbridge-perf-moe-optimization-workflow, .github/skills/nemo-mbridge-perf-moe-optimization-workflow and .opencode/skills/nemo-mbridge-perf-moe-optimization-workflow in your project.

What does Nemo Mbridge Perf Moe Optimization Workflow need to run?

SKILL.md names no scripts, command-line tools or credentials: Nemo Mbridge Perf Moe Optimization Workflow is instructions for the agent only. Our summary lists: Python 3.

Does Nemo Mbridge Perf Moe Optimization Workflow access the network?

SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Moe Optimization Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Moe Optimization Workflow use?

Nemo Mbridge Perf Moe Optimization Workflow is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Moe Optimization Workflow use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Moe Optimization Workflow?

Skills that share tags, products or a category with Nemo Mbridge Perf Moe Optimization Workflow: Tilelang Skill (slowlyC/agent-gpu-skills, 169 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Cuda Profiling (mohitmishra786/low-level-dev-skills, 253 stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Moe Optimization Workflow?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.