Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-activation-recompute --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-activation-recompute .claude/skills/nemo-mbridge-perf-activation-recompute && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-activation-recompute" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-activation-recompute into .claude/skills/nemo-mbridge-perf-activation-recompute/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-activation-recompute", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-activation-recomputeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-activation-recompute --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-activation-recompute .agents/skills/nemo-mbridge-perf-activation-recompute && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-activation-recompute" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-activation-recompute into .agents/skills/nemo-mbridge-perf-activation-recompute/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-activation-recompute", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-activation-recompute --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-activation-recompute .cursor/skills/nemo-mbridge-perf-activation-recompute && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-activation-recompute" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-activation-recompute into .cursor/skills/nemo-mbridge-perf-activation-recompute/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-activation-recompute", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-activation-recompute--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-activation-recompute --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-activation-recompute .gemini/skills/nemo-mbridge-perf-activation-recompute && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-activation-recompute" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-activation-recompute into .gemini/skills/nemo-mbridge-perf-activation-recompute/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-activation-recompute", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-activation-recomputeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-activation-recompute .github/skills/nemo-mbridge-perf-activation-recompute && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-activation-recompute" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-activation-recompute into .github/skills/nemo-mbridge-perf-activation-recompute/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-activation-recompute", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-activation-recompute --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-activation-recompute .opencode/skills/nemo-mbridge-perf-activation-recompute && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-activation-recompute" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-activation-recompute into .opencode/skills/nemo-mbridge-perf-activation-recompute/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-activation-recompute", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-activation-recomputeValidate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.
Nemo Mbridge Perf Activation Recompute is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. Use for activation memory OOMs or regressions involving recomputegranularity, recomputenumlayers, recomputemodules, recomputemethod, selective recompute, full recompute, or activation checkpointing.
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are python).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
docs.nvidia.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Activation Recompute loads about 4.7k tokens when it runs. Until then it costs about 94 tokens; SKILL.md has 2,264 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 2,264 words, ~4,729 tokens.
.claude/skills/nemo-mbridge-perf-activation-recompute/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Stable docs: @docs/training/activation-recomputation.md Card: @skills/nemo-mbridge-perf-activation-recompute/card.yaml
<!-- Guidance refreshed: 2026-08-12. -->
Activation recompute (activation checkpointing) trades additional forward work during backward for lower retained-activation memory. The useful checkpoint boundary depends on the model architecture, attention backend, parallelism, and the tensor that actually drives the per-rank peak.
max_memory_allocated() with max_memory_reserved() on every rank.recompute_modules=[] is valid and useful for this comparison.core_attn is the common first candidate. It is strongest when unfused attention materializes score/probability tensors. With Transformer Engine fused or Flash Attention, compare it against [] because those backends already rematerialize attention internals.mla_up_proj when expanded Q/K/V projections dominate. Add core_attn only when the attention-core state still matters.moe_act when the expert intermediate activation dominates; add layernorm when norm outputs are material. Use whole moe recompute only after accounting for the extra expert compute and communication it replays.mlp can save the whole dense-MLP activation region, but it usually costs more compute than a narrow output-discard boundary.Megatron Core's cpu_offloading=True is an alternative when PCIe/NVLink transfer overhead is preferable to replayed compute. It cannot be combined with activation recompute and is not compatible with pipeline parallelism greater than one.
cfg.model.recompute_granularity = "selective"
cfg.model.recompute_modules = ["core_attn"] # Common standard-attention candidate, not a universal default.Use the decision table below to replace or extend that list for MLA, MoE, dense-MLP, or GDN workloads.
cfg.model.recompute_granularity = "full"
cfg.model.recompute_method = "uniform"
cfg.model.recompute_num_layers = 1uniform: checkpoint fixed groups of recompute_num_layers transformer layers.block: checkpoint the first recompute_num_layers layers on each pipeline stage, with virtual-pipeline-aware distribution.The currently pinned Megatron Core accepts these labels. A development branch can add model-specific labels, so validate against the exact target revision rather than copying a list across branches.
| Module | Checkpoint boundary | When to test it | Main cost or caveat |
|---|---|---|---|
core_attn | Core attention | Standard attention, especially an unfused backend retaining attention intermediates | Replays attention. Incremental savings can be small with TE fused/Flash Attention; context parallelism can replay attention communication. |
mla_up_proj | MLA Q/KV up-projection plus RoPE region | MLA models retaining expanded Q/K/V tensors | Replays the MLA expansion path. It is a distinct, potentially additive boundary from core_attn. |
layernorm | Input and pre-MLP normalization outputs | Norm outputs contribute materially to the peak, often alongside MoE or MLA boundaries | Usually narrow, but savings depend on hidden size, sequence length, and which graph paths are active. |
moe_act | Activation output between grouped expert FC1 and FC2 | Grouped MoE expert-intermediate activations dominate | Narrow output-discard checkpoint. It does not replay dispatch, FC1, or FC2, but has FP8 delayed-scaling restrictions. |
mlp | Whole dense MLP | Dense layers dominate after narrower boundaries are exhausted | Replays the complete dense MLP. It has no effect on layers whose MLP is MoE. |
moe | Whole MoE forward | A broad MoE region must be discarded to make the workload fit | Replays routing, dispatch/combine communication, experts, and shared-expert work. It is incompatible with expert-parallel overlap. |
shared_experts | Non-overlapped shared-expert MLP | Shared experts are a distinct material peak | Replays the shared-expert MLP and is invalid with shared-expert overlap. Outer moe already removes its original-forward saves, but nesting can still change the transient backward-replay peak. |
gdn_norm_out | GDN gated-normalization output | GDN/hybrid models retain this output | Replays the normalization and its HP-to-CP all-to-all path. |
For example, DeepSeek V4 configurations can use the model-specific mhc
label only with their required Megatron Core development branch. It is not a
portable label for the pinned revision and therefore is not included in the
table above.
Common performance configurations consequently fall into several patterns rather than one universal list:
core_attn;mla_up_proj, sometimes with mlp;moe_act or layernorm plus moe_act;moe plus layernorm.These are candidate patterns, not an ordering guarantee. Peak attribution and matched measurements decide the final list.
For every candidate, capture:
recompute_granularity, module list, method, and layer count;max_memory_allocated() and max_memory_reserved();Use a matched no-recompute control and change one recompute choice at a time. Peak memory from different jobs, backends, or parallel layouts is not a module-ranking benchmark.
Do not call a candidate successful merely because it advances farther than the control. Run through optimizer-state initialization and multiple steady-state steps: selective recompute can move the memory wall from forward into gradient synchronization or the optimizer without making the workload viable.
A 2026-08-12 short-run study used the exact Bridge revision
600d069b824dd5ce50367a311a5a3244478faf22 and Megatron Core revision
24bad8e677d22625d86ef2a54c9506b6e4992c93. The Moonlight 16B BF16
pretraining recipe ran on 8 H100 80GB GPUs with sequence length 4096, MBS=1,
GBS=4, TP=2, PP=1, CP=1, EP=8, mock data, and 20 steps. This model mixes one
dense layer with 26 MLA+MoE layers. Each row changed only
recompute_modules; all 20 losses were finite with zero skipped or NaN
iterations.
Peak allocated memory is the maximum post-optimizer value reported after iteration 2. Time and throughput are means over iterations 11--20.
| Selective modules | Peak allocated (GB) | Step time (ms) | TFLOP/s/GPU | Allocated vs [] | Time vs [] |
|---|---|---|---|---|---|
[] | 36.618 | 457.18 | 77.50 | control | control |
core_attn | 36.614 | 474.80 | 74.44 | -0.01% | +3.85% |
mla_up_proj | 35.902 | 480.89 | 73.72 | -1.96% | +5.19% |
mla_up_proj, mlp | 35.917 | 496.73 | 71.50 | -1.91% | +8.65% |
moe_act | 35.941 | 466.26 | 75.50 | -1.85% | +1.99% |
layernorm, moe_act | 35.949 | 506.53 | 70.27 | -1.83% | +10.79% |
For this exact workload, moe_act is the best first boundary: it recovered
nearly as much allocated memory as mla_up_proj for less replay cost.
mla_up_proj is the next candidate if its roughly 39 MB additional reduction
matters. Adding mlp to mla_up_proj or layernorm to moe_act did not
improve the observed peak and made steps slower. Explicit core_attn added
cost without material memory benefit under fused attention.
Maximum reserved memory stayed near 40 GB and did not fall monotonically. That is allocator caching, not contrary evidence: boundary selection in this study is based on allocated memory and successful end-to-end steps.
The same 2026-08-12 study used the native 16-H100 BF16 performance recipe for
the 52-layer hybrid Mamba/fused-attention MoE model. The matched short-run
configuration used sequence length 8192, MBS=1, GBS=16, TP=1, PP=1, CP=1,
EP=8, DP=16, expert-DP=2, HybridEP, grouped GEMM, TE CUDA graphs for attention and Mamba,
mock data, and 12 steps. Each row changed only recompute_modules.
| Selective modules | Outcome | Rank-0 measured peak | Failure or steady-state evidence |
|---|---|---|---|
[] | OOM after iteration 1 | 66.297 GB after iteration 1 | Iteration-2 MoE router allocation failed; hot ranks had about 72.9 GiB allocated. |
core_attn | OOM in iteration 1 | not comparable | Grouped-expert linear allocation failed; explicit attention recompute did not make the fused-attention workload fit. |
moe_act | OOM after iteration 1 | 62.103 GB after iteration 1 | 4.194 GB (6.33%) below the control at the matched checkpoint, but the iteration-2 output projection still needed 2 GiB. |
layernorm, moe_act | OOM in iteration 1 | not comparable | Output projection still needed 2 GiB; CUDA-graph private pools were material. |
moe | completed 12 steps | 64.653 GB after iteration 2 | 657.42 ms and 277.72 TFLOP/s/GPU over iterations 7--12. |
moe, layernorm | completed 12 steps | 63.639 GB after iteration 2 | 677.62 ms and 270.62 TFLOP/s/GPU over iterations 7--12. |
Both successful rows had finite losses and zero skipped or NaN iterations.
For this exact capacity-limited recipe, whole-moe recompute is the smallest
tested passing boundary. Adding layernorm recovered another 1.014 GB (1.57%)
of rank-0 peak at 3.07% higher step time, so the recipe's broader combination
is justified when that headroom is required. Narrow moe_act produced real
activation relief but did not make the whole training step viable.
An exploratory native 8-H100 layout failed during FP32 optimizer-state initialization even at sequence length 4096. That is optimizer capacity, not a selective-boundary throughput baseline; no timing comparison from those runs is used here.
These measurements do not define one ranking. Moonlight fit with an empty
control and favored narrow moe_act; Nemotron required broad whole-moe
recompute; historical dense Llama evidence found whole-mlp replay costly and
lacked an empty control. The correct first candidate is therefore the narrowest
boundary implicated by the architecture and peak, followed by broader replay
only when the narrow choice does not pass the complete step.
recompute_granularity="selective" uses recompute_modules; an empty list is accepted as an explicit control.recompute_granularity="full" uses recompute_method and recompute_num_layers; selective labels do not apply.TransformerConfig validator as the source of truth.core_attn may still change retained inputs/outputs, but it must earn its place in a matched [] comparison.moe recompute is incompatible with expert-parallel overlap because backward replay would repeat the overlapped routing/communication region.shared_experts recompute is incompatible with shared-expert overlap.moe_act applies to grouped-GEMM experts and is the narrower choice when only the expert activation needs to be discarded.mlp targets dense MLPs and is a no-op on MoE layers; mixed dense/MoE models can still benefit on their dense layers.moe_act and layernorm recompute are not supported with FP8 delayed scaling and require a compatible Transformer Engine version.mla_up_proj.cuda_graph_impl="full_iteration" in the pinned Megatron Core. Otherwise disable CUDA graphs; scoped/local graph capture is not a substitute for full-iteration capture here.Historical H100 measurements from Bridge PR #3107 used Llama 3 70B SFT on 32 H100 80GB GPUs with FP8 current scaling, sequence length 4096, micro-batch size 1, global batch size 32, TP=4, PP=4, VPP=5, and DP=2:
| Configuration | TFLOP/s/GPU | Peak memory |
|---|---|---|
core_attn baseline in that run | ~704 | 58.8 GB (OOM on rank 0) |
mlp | 593.6 | 55.6 GB |
mlp + core_attn | 586.8 | 55.6 GB |
core_attn + layernorm | ~702 | 59.6 GB (OOM on rank 0) |
| Golden throughput recorded in the PR context | 709.93 | Not a paired memory measurement |
Limitations of this evidence:
3rdparty/Megatron-LM/megatron/core/transformer/transformer_config.py3rdparty/Megatron-LM/megatron/core/tensor_parallel/random.py3rdparty/Megatron-LM/megatron/core/transformer/attention.py3rdparty/Megatron-LM/megatron/core/transformer/multi_latent_attention.py3rdparty/Megatron-LM/megatron/core/transformer/transformer_layer.py3rdparty/Megatron-LM/megatron/core/transformer/moe/experts.py3rdparty/Megatron-LM/megatron/core/transformer/moe/moe_layer.py3rdparty/Megatron-LM/megatron/core/ssm/gated_delta_net/gdn.py| Symptom | Likely cause | Next action |
|---|---|---|
core_attn gives little or no peak reduction | Fused/Flash attention already rematerializes the expensive internals, or the peak is elsewhere | Compare with [], attribute the peak, then test the architecture-specific boundary such as mla_up_proj or moe_act. |
MLA still OOMs after core_attn | Expanded Q/K/V projection tensors, not attention-core tensors, dominate | Test mla_up_proj; add core_attn only if matched evidence supports it. |
| MoE peak remains high | Expert intermediate or norm outputs dominate | Test moe_act, then layernorm; reserve whole moe for broader pressure. |
| Expert-overlap validation fails | Whole-moe or shared_experts recompute conflicts with overlap | Keep overlap and use a compatible inner boundary, or disable overlap and remeasure the entire configuration. |
| A selected label has no measurable effect | That module is absent or inactive on the measured layers, or graph capture bypassed the wrapper | Inspect the provider/layer mix and final graph scope; for example, mlp is ineffective on pure-MoE layers. |
| Full recompute plus CUDA graphs asserts | Graph implementation is not full-iteration | Set cuda_graph_impl="full_iteration" or disable CUDA graphs. |
| Reserved memory is high but allocated memory is stable | Allocator fragmentation or caching | Try PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True before adding recompute. |
| OOM moves to a different rank after enabling recompute | Pipeline/virtual-pipeline layer distribution changed the bottleneck | Compare per-rank peaks and tune full block/uniform placement or selective boundaries for the actual hot stage. |
| A candidate gets farther but still OOMs | Recompute moved the peak into gradient synchronization or optimizer-state initialization | Record the changed failure stage as diagnostic evidence, but require optimizer initialization and multiple steady steps before calling it a pass. |
docs/performance-guide.mdskills/nemo-mbridge-perf-memory-tuning/SKILL.mdskills/nemo-mbridge-perf-cuda-graphs/SKILL.mdskills/nemo-mbridge-perf-cpu-offloading/SKILL.md© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-activation-recompute of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Nemo Mbridge Perf Activation Recompute next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Activation Recompute this skillNVIDIA/skills | 3.6k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.8k | Automated safety check: Pass | None | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
LMIXR/CV_Deployment_skill
基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute. Nemo Mbridge Perf Activation Recompute is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use selective and full activation recompute in Megatron Bridge to reduce GPU memory usage at the cost of extra compute.
Nemo Mbridge Perf Activation Recompute fits situations like: activation memory OOMs; regressions involving recomputegranularity; recomputenumlayers; recomputemodules.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-activation-recompute in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-activation-recompute in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-activation-recompute in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-activation-recompute in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-activation-recompute -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-activation-recompute, .gemini/skills/nemo-mbridge-perf-activation-recompute, .github/skills/nemo-mbridge-perf-activation-recompute and .opencode/skills/nemo-mbridge-perf-activation-recompute in your project.
SKILL.md names no scripts, command-line tools or credentials: Nemo Mbridge Perf Activation Recompute is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: docs.nvidia.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Activation Recompute is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Activation Recompute: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 938 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.