Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cpu-offloading --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-cpu-offloading .claude/skills/nemo-mbridge-perf-cpu-offloading && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-cpu-offloading" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading into .claude/skills/nemo-mbridge-perf-cpu-offloading/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cpu-offloading", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloadingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cpu-offloading --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-cpu-offloading .agents/skills/nemo-mbridge-perf-cpu-offloading && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-cpu-offloading" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading into .agents/skills/nemo-mbridge-perf-cpu-offloading/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cpu-offloading", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cpu-offloading --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-cpu-offloading .cursor/skills/nemo-mbridge-perf-cpu-offloading && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-cpu-offloading" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading into .cursor/skills/nemo-mbridge-perf-cpu-offloading/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cpu-offloading", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-cpu-offloading--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cpu-offloading --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-cpu-offloading .gemini/skills/nemo-mbridge-perf-cpu-offloading && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-cpu-offloading" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading into .gemini/skills/nemo-mbridge-perf-cpu-offloading/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cpu-offloading", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cpu-offloadingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-cpu-offloading .github/skills/nemo-mbridge-perf-cpu-offloading && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-cpu-offloading" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading into .github/skills/nemo-mbridge-perf-cpu-offloading/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cpu-offloading", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cpu-offloading --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-cpu-offloading .opencode/skills/nemo-mbridge-perf-cpu-offloading && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-cpu-offloading" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cpu-offloading into .opencode/skills/nemo-mbridge-perf-cpu-offloading/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cpu-offloading", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-cpu-offloadingValidate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.
Nemo Mbridge Perf Cpu Offloading is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Cpu Offloading loads about 2.3k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 575 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 575 words, ~2,329 tokens.
.claude/skills/nemo-mbridge-perf-cpu-offloading/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Two independent mechanisms to move data from GPU to CPU memory:
| Mechanism | Config namespace | What gets offloaded | PP restriction |
|---|---|---|---|
| Activation offloading | model.cpu_offloading* | Activations (and optionally weights) per transformer layer | PP must be 1 |
| Optimizer offloading | optimizer.optimizer_cpu_offload | Adam optimizer states (momentum + variance) via HybridDeviceOptimizer | None |
| Situation | Recommendation |
|---|---|
| Large MoE model (30B+), needs PP > 1 | Optimizer offloading — activation offloading is blocked by PP=1 |
| Small/medium model, PP=1 fits, activation memory dominates | Activation offloading |
| Want tunable memory-speed tradeoff | Optimizer offloading with fractional optimizer_offload_fraction |
| Throughput is top priority | Don't enable — offloading always adds overhead |
| CUDA graphs are needed | Only optimizer offloading — activation offloading is incompatible |
| Memory pressure is moderate | Optimizer offload at 25–50% fraction for best efficiency |
cfg.optimizer.optimizer_cpu_offload = True
cfg.optimizer.optimizer_offload_fraction = 1.0
cfg.optimizer.overlap_cpu_optimizer_d2h_h2d = TrueCLI overrides:
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
optimizer.overlap_cpu_optimizer_d2h_h2d=Truecfg.model.cpu_offloading = True
cfg.model.cpu_offloading_num_layers = 16
cfg.model.cpu_offloading_activations = True
cfg.model.cpu_offloading_weights = False
cfg.model.pipeline_model_parallel_size = 1
cfg.model.recompute_granularity = None
cfg.model.cuda_graph_impl = "none"| Parameter | Default | Description |
|---|---|---|
optimizer_cpu_offload | False | Master switch |
optimizer_offload_fraction | 0.0 | Fraction of optimizer states on CPU (0.0–1.0) |
overlap_cpu_optimizer_d2h_h2d | False | Overlap GPU↔CPU transfers with compute |
use_torch_optimizer_for_cpu_offload | False | Use torch.optim instead of fused optimizer for CPU portion |
| Parameter | Default | Description |
|---|---|---|
cpu_offloading | False | Master switch |
cpu_offloading_num_layers | 0 | Number of transformer layers to offload (0 to num_layers-1) |
cpu_offloading_activations | True | Offload activations |
cpu_offloading_weights | False | Offload weights |
cpu_offloading_double_buffering | False | Double-buffer across layers while reloading |
pipeline_model_parallel_size must be 1recompute_granularity must be Nonefine_grained_activation_offloadingcpu_offloading_num_layers must be in [0, num_layers-1)use_distributed_optimizer = True (default in most recipes)optimizer_offload_fraction must be in [0.0, 1.0]Activation offloading is blocked for Qwen3-30B-A3B and similar large MoE models. The PP=1 constraint means each GPU holds all 48 layers; model weights + optimizer states alone (~70 GB) exceed H100 80 GB capacity.
uv run python scripts/training/run_recipe.py \
--recipe qwen3_30b_a3b_pretrain_config \
optimizer.optimizer_cpu_offload=True \
optimizer.optimizer_offload_fraction=0.5 \
train.train_iters=20 \
train.global_batch_size=8 \
train.micro_batch_size=1uv run python -m pytest \
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cpu_offload" \
tests/unit_tests/peft/test_utils.py -k "cpu_offload" -q if self.cpu_offloading and (
self.cpu_offloading_num_layers < 0 or self.cpu_offloading_num_layers >= self.num_layers
):
raise ValueError(...)
if self.cpu_offloading and self.pipeline_model_parallel_size > 1:
raise ValueError(
"Currently there is no support for Pipeline parallelism with CPU offloading"
)
if self.cpu_offloading and self.recompute_granularity is not None:
raise ValueError(
"CPU offloading does not work when activation recomputation is enabled"
) if self.cpu_offloading:
raise ValueError("CUDA graphs not supported with CPU offloading.") if self.fine_grained_activation_offloading:
assert (
not self.cpu_offloading
), "fine_grained_activation_offloading cannot be enabled with cpu_offloading." if config.optimizer_cpu_offload:
# ... setup cpu/gpu optimizer classes ...
optimizer = HybridDeviceOptimizer(
param_groups,
offload_fraction=config.optimizer_offload_fraction,
cpu_optimizer_cls=cpu_optimizer_cls,
gpu_optimizer_cls=gpu_optimizer_cls,
overlap_cpu_optimizer_d2h_h2d=config.overlap_cpu_optimizer_d2h_h2d,
pin_cpu_grads=config.pin_cpu_grads,
pin_cpu_params=config.pin_cpu_params,
) assert not config.cpu_offloading and config.recompute_granularity is None, "Cudagraphs not supported" if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_in(x)
x = self.activation(x)
if self.config.cpu_offloading and self.config.cpu_offloading_activations:
x.activation_offloading = True
x, _ = self.linear_out(x)| Symptom | Likely Cause | How To Confirm | Fix |
|---|---|---|---|
Currently there is no support for Pipeline parallelism with CPU offloading | Activation offload + PP > 1 | Check pipeline_model_parallel_size | Set PP=1 or use optimizer offloading |
CPU offloading does not work when activation recomputation is enabled | Activation offload + recompute | Check recompute_granularity | Set recompute_granularity=null |
fine_grained_activation_offloading cannot be enabled with cpu_offloading | Both offloading modes enabled | Check both flags | Use one or the other |
CUDA graphs not supported with CPU offloading | CUDA graphs + activation offload | Check cuda_graph_impl | Set cuda_graph_impl="none" |
| OOM with activation offloading | Model too large for PP=1 | Check allocated memory vs 80 GB | Use optimizer offloading with PP > 1 |
| Extreme slowdown (>4x) | 100% optimizer offload, CPU Adam bottleneck | Compare iter time at different fractions | Reduce fraction or enable overlap_cpu_optimizer_d2h_h2d |
| OOM at partial optimizer offload | Insufficient offload for this config | Check memory at different fractions | Increase fraction or add PP |
fine_grained_activation_offloading is a separate module-level approach
that works with PP > 1 but cannot be combined with layer-level
cpu_offloading.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-cpu-offloading of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Nemo Mbridge Perf Cpu Offloading next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Cpu Offloading this skillNVIDIA/skills | 3.5k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~2.8k | Automated safety check: Pass | None | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 144 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
LMIXR/CV_Deployment_skill
基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer. Nemo Mbridge Perf Cpu Offloading is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use CPU offloading in Megatron Bridge, including layer-level activation offloading and fractional optimizer state offloading with HybridDeviceOptimizer.
Nemo Mbridge Perf Cpu Offloading fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-cpu-offloading in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-cpu-offloading in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-cpu-offloading in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-cpu-offloading in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cpu-offloading -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-cpu-offloading, .gemini/skills/nemo-mbridge-perf-cpu-offloading, .github/skills/nemo-mbridge-perf-cpu-offloading and .opencode/skills/nemo-mbridge-perf-cpu-offloading in your project.
Going by SKILL.md and its folder, Nemo Mbridge Perf Cpu Offloading needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Cpu Offloading is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Cpu Offloading: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.