Official agent skill

Nemo Mbridge Perf Memory Tuning

by NVIDIA in NVIDIA/skills

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM…

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Perf Memory Tuning

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-memory-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-memory-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-memory-tuning .claude/skills/nemo-mbridge-perf-memory-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-memory-tuning
GitHub stars
3.5k
Token cost
~3.6k tokens
SKILL.md length
1,342 words
Files
6
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM…

  • Works in 7 steps: Set… → For LoRA with sequence parallelism,… → Add selective activation recompute… → …
  • Tasks that involve Fine-tuning
  • SKILL.md covers What It Is, Quick Decision, Enablement and A Note on VPP, plus 6 more sections
  • Calls uv

What it does

Nemo Mbridge Perf Memory Tuning is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering, covering Fine-tuning and GPU and accelerator computing. It works with NVIDIA AI Platform, CUDA and PyTorch. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Fine-tuning
  • Tasks that involve GPU and accelerator computing

Example prompts

  • “Use the nemo-mbridge-perf-memory-tuning skill to technique for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input…”
  • “/nemo-mbridge-perf-memory-tuning”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True first. This fixes
  2. For LoRA with sequence parallelism, enable input re-gather
  3. Add selective activation recompute (recompute_modules=[core_attn]) if
  4. Avoid increasing TP as a memory fix — doubling TP dramatically increases
  5. Avoid increasing PP at the cost of DP — halving DP doubles gradient
  6. Consider mlp recompute if still OOM. Saves ~3 GB but costs ~16% GPU
  7. CPU offloading is blocked when PP > 1.

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Memory Tuning loads about 3.6k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 1,342 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,342 words, ~3,604 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-memory-tuning/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-memory-tuning
description
Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.
license
Apache-2.0
when_to_use
GPU OOM errors, reducing peak memory, reducing LoRA or PEFT activation memory with sequence parallelism, or tracing an OOM regression to a specific commit or…

Memory Tuning

Stable docs: @docs/parallelisms.md Card: @skills/nemo-mbridge-perf-memory-tuning/card.yaml

What It Is

GPU OOM failures during training often stem from memory fragmentation rather than raw capacity. PyTorch's default CUDA allocator can leave unusable gaps between allocations. The single most effective fix is:

bash
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

This tells PyTorch to use expandable (non-fixed-size) memory segments, which dramatically reduces fragmentation and often eliminates borderline OOM without any model or parallelism changes.

Beyond fragmentation, actual peak memory is determined by:

  • Parameter + optimizer state memory — controlled by TP, PP, DP sharding (distributed optimizer, FSDP)
  • Activation memory — controlled by activation recompute, sequence length, micro-batch size, and PEFT-specific retention of gathered inputs
  • Temporary / workspace memory — CUDA kernels, NCCL buffers, CUDA graphs

For configuration planning, use the Bridge theoretical estimator before launching large jobs:

python
from megatron.bridge.training.utils.theoretical_memory_utils import estimate_training_memory

estimate = estimate_training_memory(cfg, num_microbatches=num_microbatches)

The estimator reports the most-loaded GPU shard and separates dense/embedding, routed MoE expert, and activation components. It does not include allocator fragmentation, CUDA/NCCL workspace, CUDA graph buffers, token imbalance, or dispatcher workspace, so validate final configs with runtime memory metrics.

Quick Decision

When a training run OOMs or is close to the memory limit:

  1. Set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True first. This fixes fragmentation-induced OOM with zero performance cost. Most Slurm launch templates already include it.
  2. For LoRA with sequence parallelism, enable input re-gather (LoRA(sequence_parallel_input_regather=True)). This avoids retaining the full gathered LoRA-A input in every eligible layer; it has no effect when SP is disabled.
  3. Add selective activation recompute (recompute_modules=[core_attn]) if not already enabled. See @skills/nemo-mbridge-perf-activation-recompute/SKILL.md.
  4. Avoid increasing TP as a memory fix — doubling TP dramatically increases NVLink all-reduce volume and often kills throughput (-28% on Llama3 70B).
  5. Avoid increasing PP at the cost of DP — halving DP doubles gradient accumulation steps and hurts throughput (~6%).
  6. Consider mlp recompute if still OOM. Saves ~3 GB but costs ~16% GPU utilization on large dense models (Llama3 70B).
  7. CPU offloading is blocked when PP > 1.

Enablement

Set in the job's environment before launching:

bash
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

In Slurm scripts this is typically placed alongside other env vars:

bash
export CUDA_DEVICE_MAX_CONNECTIONS=1
export NVTE_ALLOW_NONDETERMINISTIC_ALGO=1
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

No model config changes needed. Zero throughput cost.

Parallelism resizing

If the model genuinely does not fit (not fragmentation), adjust parallelism:

StrategyMemory effectThroughput costNotes
Increase PP (keeping DP)Fewer layers per stageModerate (~6% if DP halved)Only if GPU count allows
Increase TPFewer params per GPUSevere (-28% on 70B)Last resort
Distributed optimizerShards optimizer state across DP ranks~1-2%Recommended for large models
FSDPShards params + grads + optimizerVariesSee @skills/nemo-mbridge-perf-megatron-fsdp/SKILL.md
Activation recompute

See @skills/nemo-mbridge-perf-activation-recompute/SKILL.md for full details.

PEFT + sequence-parallel input re-gather

For LoRA training with sequence parallelism, eligible column-parallel linear_qkv and linear_fc1 adapters consume a gathered LayerNorm output. Because LoRA-A is trainable, the default path retains that full gathered input until backward for the LoRA-A weight gradient.

Enable input re-gather when constructing the PEFT config:

python
from megatron.bridge.peft.lora import LoRA

cfg.peft = LoRA(
    # Keep the recipe's existing LoRA settings here.
    sequence_parallel_input_regather=True,
)

With this option, forward still materializes the full input temporarily for the LoRA-A GEMM, but MCore autograd retains only its sequence-local shard. Backward asynchronously gathers the full input again, overlaps the collective with dgrad when possible, computes the LoRA-A weight gradient, and then reuses the temporary communication buffer.

This is a memory-for-communication tradeoff, not conventional activation checkpointing: no LayerNorm, attention, MLP, or LoRA GEMM is rerun. Some throughput degradation is expected, and the benefit grows with the amount of eligible LoRA-A activation retained. The option has no effect when sequence parallelism is disabled.

CPU offloading
python
cfg.model.cpu_offloading = True

Incompatible with PP > 1. Only usable when pipeline_model_parallel_size = 1.

A Note on VPP

Virtual pipeline parallelism (VPP) is primarily a throughput optimization that reduces pipeline bubble overhead by interleaving smaller model chunks. Its effect on peak memory is minimal — changing VPP does not meaningfully change the total activation, parameter, or optimizer memory on a GPU.

In earlier experiments we incorrectly attributed an OOM fix to VPP tuning (VPP 5→10). The actual fix was PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True which eliminated memory fragmentation. The VPP=10 run actually used slightly more peak memory (60.2 GB vs 58.8 GB) but did not OOM because expandable segments prevented fragmentation.

VPP should be tuned for pipeline bubble reduction (see @docs/parallelisms.md), not as a memory fix.

Compatibility and Constraints

  • expandable_segments:True is incompatible with --use-nccl-ub (NCCL user-buffer registration). See Megatron-FSDP docs.
  • When using CUDA graphs with expandable_segments:True, set NCCL_GRAPH_REGISTER=0 (required on pre-Blackwell GPUs, enforced by MCore CudaGraphManager).
  • CPU offloading requires pipeline_model_parallel_size = 1.
  • Distributed optimizer requires use_distributed_optimizer = True in the optimizer config.
  • sequence_parallel_input_regather applies only to eligible non-expert column-parallel LoRA-A projections. Row-parallel adapters, expert adapters, TP=1, CUDA graphs, CPU activation offload, and overlapping full-layer or selective MLP activation recompute fall back to the existing path.

Measured Results

Llama3 70B SFT on 32x H100 80GB, FP8 (Current Scaling):

  • Baseline: TP=4, PP=4, VPP=5, DP=2, MBS=1, GBS=32, seq_len=4096
  • Golden GPU utilization: 709.93 TFLOP/s/GPU
  • Regression threshold: 5%
Show full SKILL.md (561 more words)Show less
Strategy comparison: parallelism changes for memory reduction
ExperimentTPPPVPPDPTFLOP/s/GPUvs GoldenPeak Mem (GB)Result
Baseline4452~704-0.8%58.8OOM (fragmentation)
More PP4851668.0-5.9%53.2Borderline perf
More TP8451508.7-28.4%50.2Severe regression
Baseline + expandable_segments4452~704-0.8%~59Passed

Key takeaways:

  • expandable_segments:True is the winner. The baseline OOM was caused by memory fragmentation, not insufficient capacity. Setting this env var eliminated the OOM with zero throughput cost and no parallelism changes.
  • PP=8 works for memory but loses DP (2→1), meaning 32 gradient accumulation steps per batch, which hurts throughput by ~6%.
  • TP=8 is catastrophic (-28%) because doubling TP increases all-reduce communication volume proportionally across NVLink, and DP=1 means no micro-batch overlap.
CPU offloading: blocked
Experimentoffload_layersResult
Exp 42Incompatible (PP > 1)
Exp 54Incompatible (PP > 1)
Exp 66Incompatible (PP > 1)

ValueError: Currently there is no support for Pipeline parallelism with CPU offloading. This approach is blocked for any model using PP > 1.

Activation recompute: expensive alternative

Selective activation recompute with mlp saved ~3 GB peak memory but cost ~16% GPU utilization on this workload. See @skills/nemo-mbridge-perf-activation-recompute/SKILL.md for full results.

LoRA + SP input re-gather

Real-checkpoint H100 training with SQuAD showed lower peak memory in all tested configurations, with workload-dependent throughput cost:

Model/configBaseline peakInput re-gather peakMemory savedThroughput change
Qwen3-8B, TP2, seq 819247.545 GB42.814 GB4.731 GB (10.0%)-6.74%
Qwen3-30B-A3B, TP4/EP429.890 GB28.321 GB1.569 GB (5.2%)-2.89%
GPT-OSS-120B, TP2/EP852.185 GB51.371 GB0.814 GB (1.6%)-0.34%

All runs had finite losses with zero skipped or NaN iterations. Two-rank BF16 and FP32 checks matched the baseline for outputs, input gradients, LoRA-A and LoRA-B gradients, and two-microbatch fused main_grad accumulation.

Code Anchors

LoRA sequence-parallel input re-gather
text
src/megatron/bridge/peft/lora.py
    LoRA.sequence_parallel_input_regather

src/megatron/bridge/peft/utils.py
    ParallelLinearAdapter._sequence_parallel_input_regather_eligibility()
    ParallelLinearAdapter.forward()
CPU offloading PP incompatibility (MCore)
130313063rdpartyMega
        if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
            raise ValueError(
                "Currently there is no support for Pipeline parallelism with CPU offloading"
            )
VPP config and layer divisibility validation (MCore)
158115923rdpartyMega
            if pipeline_parallel_size and self.virtual_pipeline_model_parallel_size is not None:
                num_layers_per_middle_pipeline_rank = num_layers // pipeline_parallel_size
                if (
                    not num_layers_per_middle_pipeline_rank
                    % self.virtual_pipeline_model_parallel_size
                    == 0
                ):
                    raise ValueError(
                        f"number of layers on each middle pipeline rank:"
                        f"{num_layers_per_middle_pipeline_rank} must be divisible by virtual"
                        f"pipeline parallel degree {self.virtual_pipeline_model_parallel_size}"
                    )
Parallelism docs on interleaved pipeline schedule
116124docsparallelis
To minimize the pipeline bubble, the computation on each GPU can be divided into multiple subsets of layers (referred to as model chunks), rather than a single contiguous block. Enable this by setting `virtual_pipeline_model_parallel_size`:

model_config = GPTModelProvider(
    pipeline_model_parallel_size=4,
    virtual_pipeline_model_parallel_size=2,  # 2 model chunks per pipeline stage
    # ... other model parameters
)

Failure Diagnosis

SymptomCauseConfirmFix
OOM on a single rank despite headroom on othersMemory fragmentationcheck if expandable_segments:True is setset PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
OOM with expandable_segments already setGenuine capacity limitcheck nvidia-smi for param/optimizer memoryincrease PP, use distributed optimizer, or add recompute
Estimated memory exceeds GPU capacity before launchmodel state or activations genuinely too largerun estimate_training_memory and inspect the largest componentadjust PP/TP/CP/EP, distributed optimizer, or recompute before launching
LoRA + SP retains unexpectedly high activation memoryfull gathered LoRA-A inputs are retained until backwardcheck whether cfg.peft.sequence_parallel_input_regather is enabled and the target is eligibleset LoRA(sequence_parallel_input_regather=True); verify fallback constraints
ValueError: PP + CPU offloadingusing cpu_offloading with PP > 1check PP configdisable CPU offloading or set PP=1
RuntimeError with --use-nccl-ub + expandable segmentsNCCL UB incompatible with expandable allocatorcheck env varsremove expandable_segments:True or disable --use-nccl-ub

Known Limitations

  • CPU offloading is blocked when PP > 1
  • Parallelism resizing (TP/PP) often has significant throughput costs
  • The theoretical estimator is formula-based and does not replace runtime profiling or CUDA memory reports
  • LoRA input re-gather does not cover row-parallel or expert adapters and may have negligible benefit when few eligible LoRA-A activations dominate memory

Verification

Quick check that expandable_segments:True is active:

python
import os
assert "expandable_segments:True" in os.environ.get("PYTORCH_CUDA_ALLOC_CONF", "")

For Slurm jobs, verify the env var is exported before the training command in the launch script.

For LoRA + SP input re-gather, run the focused configuration tests and the real two-rank MCore backward-parity test:

bash
uv run python -m pytest \
  tests/unit_tests/peft/test_utils.py -k "sequence_parallel_input_regather" \
  tests/unit_tests/peft/test_lora.py -k "sequence_parallel_input_regather"

uv run python -m torch.distributed.run --nproc_per_node=2 -m pytest \
  tests/unit_tests/peft/test_lora_sp_input_regather_distributed.py

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-memory-tuning of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Nemo Mbridge Perf Memory Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Memory Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Memory Tuning this skillNVIDIA/skills3.5k—~3.6kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Hyperpod Version Checkerawslabs/agent-plugins9151 repos~910Automated safety check: PassApache-2.0
Spark Environment Setupwshobson/agents40k—~2kAutomated safety check: PassMIT
GPU OptimizerMathews-Tom/armory328—~3.5kAutomated safety check: NotesMIT
OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs13k1 repos~3.7kAutomated safety check: PassMIT

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    915 GitHub starsUsed in 1 repo~910 tokens
    AI & LLM EngineeringAuto-check passed
  • Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

    40k GitHub stars~2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    328 GitHub stars~3.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes
  • OpenVLA-OFT Fine-Tuning

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

    13k GitHub starsUsed in 1 repo~3.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemo Mbridge Perf Memory Tuning

What does Nemo Mbridge Perf Memory Tuning do?

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM…. Nemo Mbridge Perf Memory Tuning is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.

When should I use Nemo Mbridge Perf Memory Tuning?

Nemo Mbridge Perf Memory Tuning fits situations like: tasks that involve Fine-tuning; tasks that involve GPU and accelerator computing.

How do I install Nemo Mbridge Perf Memory Tuning in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-memory-tuning -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-memory-tuning in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-memory-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Memory Tuning in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-memory-tuning -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-memory-tuning in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-memory-tuning in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Memory Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-memory-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-memory-tuning, .gemini/skills/nemo-mbridge-perf-memory-tuning, .github/skills/nemo-mbridge-perf-memory-tuning and .opencode/skills/nemo-mbridge-perf-memory-tuning in your project.

What does Nemo Mbridge Perf Memory Tuning need to run?

Going by SKILL.md and its folder, Nemo Mbridge Perf Memory Tuning needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Perf Memory Tuning access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Memory Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Memory Tuning use?

Nemo Mbridge Perf Memory Tuning is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Memory Tuning use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Memory Tuning?

Skills that share tags, products or a category with Nemo Mbridge Perf Memory Tuning: Graphsignal (graphsignal/graphsignal, 257 stars), Hyperpod Version Checker (awslabs/agent-plugins, 915 stars), Spark Environment Setup (wshobson/agents, 40k stars) and GPU Optimizer (Mathews-Tom/armory, 328 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Memory Tuning?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.