Official agent skill

Nemo Mbridge Perf Cpu Offloading

by NVIDIA in NVIDIA/skills

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Perf Cpu Offloading

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cpu-offloading --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-cpu-offloading .claude/skills/nemo-mbridge-perf-cpu-offloading && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-cpu-offloading
GitHub stars
3.5k
Token cost
~2.3k tokens
SKILL.md length
575 words
Files
6
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

  • AI & LLM Engineering work in your project
  • SKILL.md covers References, What It Is, Quick Decision and Enablement, plus 7 more sections
  • Calls uv

What it does

Nemo Mbridge Perf Cpu Offloading is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/nemo-mbridge-perf-cpu-offloading”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Cpu Offloading loads about 2.3k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 575 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 575 words, ~2,329 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-cpu-offloading/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-cpu-offloading
description
Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.
license
Apache-2.0
when_to_use
Enabling CPU offload to reduce GPU memory, or investigating a commit that changed CPU offloading config and caused OOM or a crash; 'cpu_offloading'…

CPU Offloading

References

  • Stable docs: @docs/training/cpu-offloading.md
  • Structured metadata: @skills/nemo-mbridge-perf-cpu-offloading/card.yaml

What It Is

Two independent mechanisms to move data from GPU to CPU memory:

MechanismConfig namespaceWhat gets offloadedPP restriction
Activation offloadingmodel.cpu_offloading*Activations (and optionally weights) per transformer layerPP must be 1
Optimizer offloadingoptimizer.optimizer_cpu_offloadAdam optimizer states (momentum + variance) via HybridDeviceOptimizerNone

Quick Decision

SituationRecommendation
Large MoE model (30B+), needs PP > 1Optimizer offloading — activation offloading is blocked by PP=1
Small/medium model, PP=1 fits, activation memory dominatesActivation offloading
Want tunable memory-speed tradeoffOptimizer offloading with fractional optimizer_offload_fraction
Throughput is top priorityDon't enable — offloading always adds overhead
CUDA graphs are neededOnly optimizer offloading — activation offloading is incompatible
Memory pressure is moderateOptimizer offload at 25–50% fraction for best efficiency

Enablement

python
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = True

CLI overrides:

bash
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=True
Activation CPU offloading (small/medium models only)
python
cfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False

cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"

Config Parameter Reference

Optimizer offloading
ParameterDefaultDescription
optimizer_cpu_offloadFalseMaster switch
optimizer_offload_fraction0.0Fraction of optimizer states on CPU (0.0–1.0)
overlap_cpu_optimizer_d2h_h2dFalseOverlap GPU↔CPU transfers with compute
use_torch_optimizer_for_cpu_offloadFalseUse torch.optim instead of fused optimizer for CPU portion
Activation offloading
ParameterDefaultDescription
cpu_offloadingFalseMaster switch
cpu_offloading_num_layers0Number of transformer layers to offload (0 to num_layers-1)
cpu_offloading_activationsTrueOffload activations
cpu_offloading_weightsFalseOffload weights
cpu_offloading_double_bufferingFalseDouble-buffer across layers while reloading

Compatibility And Constraints

Activation offloading
  • pipeline_model_parallel_size must be 1
  • recompute_granularity must be None
  • Cannot combine with fine_grained_activation_offloading
  • Cannot combine with CUDA graphs
  • cpu_offloading_num_layers must be in [0, num_layers-1)
Optimizer offloading
  • Requires use_distributed_optimizer = True (default in most recipes)
  • No PP, recompute, or CUDA graph restrictions
  • optimizer_offload_fraction must be in [0.0, 1.0]
Practical: large MoE models

Activation offloading is blocked for Qwen3-30B-A3B and similar large MoE models. The PP=1 constraint means each GPU holds all 48 layers; model weights + optimizer states alone (~70 GB) exceed H100 80 GB capacity.

Minimal Runnable Command

bash
uv run python scripts/training/run_recipe.py \
  --recipe qwen3_30b_a3b_pretrain_config \
  optimizer.optimizer_cpu_offload=True \
  optimizer.optimizer_offload_fraction=0.5 \
  train.train_iters=20 \
  train.global_batch_size=8 \
  train.micro_batch_size=1

Verification

Unit tests
bash
uv run python -m pytest \
  tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
  tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q
Success criteria
  • Config validation passes for the selected offloading mode
  • Training completes without OOM or NCCL errors
  • Loss matches the non-offloaded baseline (max delta < 0.001)
  • Memory usage drops proportionally to offload fraction

Code Anchors

Show full SKILL.md (232 more words)Show less
MCore activation offload constraints
129613103rdpartyMega
        if self.cpu_offloading and (
            self.cpu_offloading_num_layers < 0 or self.cpu_offloading_num_layers >= self.num_layers
        ):
            raise ValueError(...)

        if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
            raise ValueError(
                "Currently there is no support for Pipeline parallelism with CPU offloading"
            )

        if self.cpu_offloading and self.recompute_granularity is not None:
            raise ValueError(
                "CPU offloading does not work when activation recomputation is enabled"
            )
MCore CUDA graph incompatibility
194319443rdpartyMega
            if self.cpu_offloading:
                raise ValueError("CUDA graphs not supported with CPU offloading.")
MCore fine-grained offloading mutual exclusion
142714303rdpartyMega
        if self.fine_grained_activation_offloading:
            assert (
                not self.cpu_offloading
            ), "fine_grained_activation_offloading cannot be enabled with cpu_offloading."
MCore HybridDeviceOptimizer instantiation
4805183rdpartyMegatr
        if config.optimizer_cpu_offload:
            # ... setup cpu/gpu optimizer classes ...
            optimizer = HybridDeviceOptimizer(
                param_groups,
                offload_fraction=config.optimizer_offload_fraction,
                cpu_optimizer_cls=cpu_optimizer_cls,
                gpu_optimizer_cls=gpu_optimizer_cls,
                overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
                pin_cpu_grads=config.pin_cpu_grads,
                pin_cpu_params=config.pin_cpu_params,
            )
Bridge CUDA graph guard
232234srcmegatronbri
        assert not config.cpu_offloading and config.recompute_granularity is None, "Cudagraphs not supported"
Bridge activation offloading in PEFT
621631srcmegatronbri
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_in(x)
        x = self.activation(x)
        if self.config.cpu_offloading and self.config.cpu_offloading_activations:
            x.activation_offloading = True
        x, _ = self.linear_out(x)

Failure Diagnosis

SymptomLikely CauseHow To ConfirmFix
Currently there is no support for Pipeline parallelism with CPU offloadingActivation offload + PP > 1Check pipeline_model_parallel_sizeSet PP=1 or use optimizer offloading
CPU offloading does not work when activation recomputation is enabledActivation offload + recomputeCheck recompute_granularitySet recompute_granularity=null
fine_grained_activation_offloading cannot be enabled with cpu_offloadingBoth offloading modes enabledCheck both flagsUse one or the other
CUDA graphs not supported with CPU offloadingCUDA graphs + activation offloadCheck cuda_graph_implSet cuda_graph_impl="none"
OOM with activation offloadingModel too large for PP=1Check allocated memory vs 80 GBUse optimizer offloading with PP > 1
Extreme slowdown (>4x)100% optimizer offload, CPU Adam bottleneckCompare iter time at different fractionsReduce fraction or enable overlap_cpu_optimizer_d2h_h2d
OOM at partial optimizer offloadInsufficient offload for this configCheck memory at different fractionsIncrease fraction or add PP

Known Limitations

  • Activation offloading requires PP=1, making it impractical for large models (30B+ MoE) that need pipeline parallelism.
  • Optimizer offloading throughput penalty scales linearly (~1.9x at 25%, ~4.2x at 100% for Qwen3-30B-A3B).
  • D2H/H2D overlap provides only ~7% speedup because CPU Adam compute is the dominant bottleneck.
  • fine_grained_activation_offloading is a separate module-level approach that works with PP > 1 but cannot be combined with layer-level cpu_offloading.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-cpu-offloading of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Nemo Mbridge Perf Cpu Offloading next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Cpu Offloading compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Cpu Offloading this skillNVIDIA/skills3.5k—~2.3kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS911—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent144—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    144 GitHub stars~5.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    146 GitHub stars~547 tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Nemo Mbridge Perf Cpu Offloading

What does Nemo Mbridge Perf Cpu Offloading do?

Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. Nemo Mbridge Perf Cpu Offloading is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.

When should I use Nemo Mbridge Perf Cpu Offloading?

Nemo Mbridge Perf Cpu Offloading fits situations like: AI & LLM Engineering work in your project.

How do I install Nemo Mbridge Perf Cpu Offloading in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-cpu-offloading in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-cpu-offloading in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Cpu Offloading in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-cpu-offloading in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-cpu-offloading in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Cpu Offloading in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-cpu-offloading, .gemini/skills/nemo-mbridge-perf-cpu-offloading, .github/skills/nemo-mbridge-perf-cpu-offloading and .opencode/skills/nemo-mbridge-perf-cpu-offloading in your project.

What does Nemo Mbridge Perf Cpu Offloading need to run?

Going by SKILL.md and its folder, Nemo Mbridge Perf Cpu Offloading needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Perf Cpu Offloading access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Cpu Offloading safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Cpu Offloading use?

Nemo Mbridge Perf Cpu Offloading is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Cpu Offloading use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Cpu Offloading?

Skills that share tags, products or a category with Nemo Mbridge Perf Cpu Offloading: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Cpu Offloading?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.