Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cuda-graphs --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-cuda-graphs .claude/skills/nemo-mbridge-perf-cuda-graphs && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-cuda-graphs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cuda-graphs into .claude/skills/nemo-mbridge-perf-cuda-graphs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cuda-graphs", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cuda-graphsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cuda-graphs --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-cuda-graphs .agents/skills/nemo-mbridge-perf-cuda-graphs && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-cuda-graphs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cuda-graphs into .agents/skills/nemo-mbridge-perf-cuda-graphs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cuda-graphs", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cuda-graphs --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-cuda-graphs .cursor/skills/nemo-mbridge-perf-cuda-graphs && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-cuda-graphs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cuda-graphs into .cursor/skills/nemo-mbridge-perf-cuda-graphs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cuda-graphs", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-cuda-graphs--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cuda-graphs --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-cuda-graphs .gemini/skills/nemo-mbridge-perf-cuda-graphs && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-cuda-graphs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cuda-graphs into .gemini/skills/nemo-mbridge-perf-cuda-graphs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cuda-graphs", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cuda-graphsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-cuda-graphs .github/skills/nemo-mbridge-perf-cuda-graphs && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-cuda-graphs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cuda-graphs into .github/skills/nemo-mbridge-perf-cuda-graphs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cuda-graphs", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-cuda-graphs --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-cuda-graphs .opencode/skills/nemo-mbridge-perf-cuda-graphs && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-cuda-graphs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-cuda-graphs into .opencode/skills/nemo-mbridge-perf-cuda-graphs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-cuda-graphs", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-cuda-graphsValidate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
Nemo Mbridge Perf Cuda Graphs is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It sits in AI & LLM Engineering, covering Deep learning. It works with CUDA and NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Cuda Graphs loads about 3.5k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 915 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 915 words, ~3,521 tokens.
.claude/skills/nemo-mbridge-perf-cuda-graphs/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Stable documentation: @docs/training/cuda-graphs.md Card: @skills/nemo-mbridge-perf-cuda-graphs/card.yaml
<!-- NVSkills CI refresh: 2026-06-15. No instruction changes. -->
CUDA graphs capture GPU operations once and replay them with minimal host-driver overhead. Bridge supports two implementations:
cuda_graph_impl | Mechanism | Scope support |
|---|---|---|
"local" | MCore FullCudaGraphWrapper wrapping entire fwd+bwd | full_iteration |
"transformer_engine" | TE make_graphed_callables() per layer | attn, mlp, moe, moe_router, moe_preprocess, mamba |
Start with TE-scoped graphs for most training workloads, then verify replay timing against eager on the same dispatcher, layout, and container:
attn, then optionally mlpattn moe_router moe_preprocessUse local + full_iteration only when you specifically want full-iteration
capture and can satisfy the tighter constraints.
For recompute-heavy workloads:
local full-iteration graphs or away
from graphs entirelyRelated docs:
cfg.model.cuda_graph_impl = "local"
cfg.model.cuda_graph_scope = ["full_iteration"]
cfg.model.cuda_graph_warmup_steps = 3
cfg.model.use_te_rng_tracker = True
cfg.rng.te_rng_tracker = True
cfg.rerun_state_machine.check_for_nan_in_loss = False
cfg.ddp.check_for_nan_in_grad = Falsecfg.model.cuda_graph_impl = "transformer_engine"
cfg.model.cuda_graph_scope = ["attn"] # or ["attn", "mlp"]
cfg.model.cuda_graph_warmup_steps = 3
cfg.model.use_te_rng_tracker = True
cfg.rng.te_rng_tracker = Truecfg.model.cuda_graph_impl = "transformer_engine"
cfg.model.cuda_graph_scope = ["attn", "moe_router", "moe_preprocess"]
cfg.model.cuda_graph_warmup_steps = 3
cfg.model.use_te_rng_tracker = True
cfg.rng.te_rng_tracker = Trueuv run python scripts/performance/run_script.py \
-m qwen \
-mr qwen3_30b_a3b \
--task pretrain \
-g h100 \
-c bf16 \
-ng 16 \
--cuda_graph_impl transformer_engine \
--cuda_graph_scope attn,moe_router,moe_preprocess \
...Valid CLI values live in scripts/performance/argument_parser.py:
VALID_CUDA_GRAPH_IMPLS: ["none", "local", "transformer_engine"]VALID_CUDA_GRAPH_SCOPES: ["full_iteration", "attn", "mlp", "moe", "moe_router", "moe_preprocess", "mamba"]The performance harness uses a comma-separated --cuda_graph_scope value and
auto-enables model.use_te_rng_tracker plus rng.te_rng_tracker when
--cuda_graph_impl is not none.
use_te_rng_tracker = True (enforced in gpt_provider.py)full_iteration scope only with cuda_graph_impl = "local"full_iteration scope requires check_for_nan_in_loss = Falsemoe scope and moe_router scopePYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, set
NCCL_GRAPH_REGISTER=0 (MCore enforces for local impl on arch < sm_100;
TE impl asserts unconditionally)moe_preprocess scope requires moe_router scope to also be set # CUDA graph scope validation: check_for_nan_in_loss must be disabled with full_iteration graph
if self.model.cuda_graph_impl == "local" and CudaGraphScope.full_iteration in self.model.cuda_graph_scope:
assert not self.rerun_state_machine.check_for_nan_in_loss, (
"check_for_nan_in_loss must be disabled when using full_iteration CUDA graph. "
"Set rerun_state_machine.check_for_nan_in_loss=False."
)
if self.model.cuda_graph_impl == "none":
self.model.cuda_graph_scope = [] if self.cuda_graph_impl != "none":
assert getattr(self, "use_te_rng_tracker", False), (
"Transformer engine's RNG tracker is required for cudagraphs, it can be "
"enabled with use_te_rng_tracker=True'." # Capture CUDA Graphs.
cuda_graph_helper = None
if model_config.cuda_graph_impl == "transformer_engine":
cuda_graph_helper = TECudaGraphHelper(...)
# ...
if config.model.cuda_graph_impl == "local" and CudaGraphScope.full_iteration in config.model.cuda_graph_scope:
forward_backward_func = FullCudaGraphWrapper(
forward_backward_func, cuda_graph_warmup_steps=config.model.cuda_graph_warmup_steps
) # Capture CUDA Graphs after warmup.
if (
model_config.cuda_graph_impl == "transformer_engine"
and cuda_graph_helper is not None
and not cuda_graph_helper.graphs_created()
and global_state.train_state.step - start_iteration == model_config.cuda_graph_warmup_steps
):
if model_config.cuda_graph_warmup_steps > 0 and should_toggle_forward_pre_hook:
disable_forward_pre_hook(model, param_sync=False)
cuda_graph_helper.create_cudagraphs()
if model_config.cuda_graph_warmup_steps > 0 and should_toggle_forward_pre_hook:
enable_forward_pre_hook(model)
cuda_graph_helper.cuda_graph_set_manual_hooks() _set_random_seed(
rng_config.seed,
rng_config.data_parallel_random_init,
rng_config.te_rng_tracker,
rng_config.inference_rng_tracker,
use_cudagraphable_rng=(model_config.cuda_graph_impl != "none"),
pg_collection=pg_collection,
) cuda_graph_scope = getattr(model_cfg, "cuda_graph_scope", []) or []
# ... scope parsing ...
if wgrad_in_graph_scope:
assert is_te_min_version("2.12.0"), ...
assert model_cfg.gradient_accumulation_fusion, ...
if attn_scope_enabled:
assert not model_cfg.add_bias_linear and not model_cfg.add_qkv_bias, ...def _set_cuda_graph_overrides(
recipe, cuda_graph_impl=None, cuda_graph_scope=None
):
# Sets impl, scope, and auto-enables te_rng_trackerdef _delete_cuda_graphs(cuda_graph_helper):
# Deletes FullCudaGraphWrapper and TE graph objects to free NCCL buffersCudaGraphManager: megatron/core/transformer/cuda_graphs.pyTECudaGraphHelper: megatron/core/transformer/cuda_graphs.pyFullCudaGraphWrapper: megatron/core/full_cuda_graph.pyCudaGraphScope enum: megatron/core/transformer/enums.pysrc/megatron/bridge/perf_recipes/deepseek/gb300/deepseek_v3.pysrc/megatron/bridge/perf_recipes/qwen/gb300/qwen3_moe.pysrc/megatron/bridge/perf_recipes/gpt_oss/gb300/gpt_oss.py| File | Coverage |
|---|---|
tests/unit_tests/training/test_config.py | full_iteration NaN-check constraint |
tests/unit_tests/training/test_comm_overlap.py | delay_wgrad + CUDA graph interaction |
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py | TE autocast with CUDA graphs |
tests/functional_tests/test_groups/recipes/test_llama_recipes_pretrain_cuda_graphs.py | End-to-end local and TE graph smoke tests |
tests/unit_tests/recipes/kimi/test_kimi_k2.py | TE + CUDA graph recipe config |
tests/unit_tests/recipes/gpt/test_gpt3_175b.py | TE + CUDA graph recipe config |
tests/unit_tests/recipes/qwen_vl/test_qwen25_vl_recipes.py | VLM CUDA graph settings |
TE RNG tracker is mandatory: Setting cuda_graph_impl without
use_te_rng_tracker=True and rng.te_rng_tracker=True will assert
in the provider.
full_iteration requires NaN checks disabled: The entire fwd+bwd is
captured, so loss-NaN checking cannot inspect intermediate values.
MoE scope restrictions: moe scope and moe_router scope are
mutually exclusive. Token-dropless MoE can only graph moe_router and
moe_preprocess, not the full expert dispatch.
Memory overhead: CUDA graphs pin all intermediate buffers for the
graph's lifetime (no memory reuse). TE scoped graphs add a few GB;
full-iteration graphs can increase peak memory by 1.5–2×. PP > 1
compounds overhead since each stage holds its own graph.
Delayed wgrad interaction: When delay_wgrad_compute=True and
attention or MoE router is in cuda_graph_scope, additional constraints
apply: TE >= 2.12.0, gradient_accumulation_fusion=True, and no
attention bias.
Variable-length sequences break graphs: Sequence lengths must be constant across steps. Use padded packed sequences if packing is needed.
Graph cleanup is required: CUDA graph objects hold NCCL buffer
references. Bridge handles this in _delete_cuda_graphs() at the end
of training, but early exits must call it explicitly.
Older GPU architectures: On GPUs with compute capability < 10.0
(pre-Blackwell), set NCCL_GRAPH_REGISTER=0 when using
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True. Enforced in MCore
CudaGraphManager (cuda_graphs.py:1428) and TECudaGraphHelper
(cuda_graphs.py:1697). The TE impl asserts unconditionally regardless
of arch.
CPU offloading incompatible: CUDA graphs cannot be used with CPU
offloading. Enforced in MCore transformer_config.py:1907.
MoE recompute + moe_router scope: MoE recompute is not supported
with moe_router CUDA graph scope when using cuda_graph_impl = "transformer_engine". Enforced in MCore transformer_config.py:1977.
Layer-level recompute requires full_iteration scope: Using
recompute_granularity="full" with recompute_num_layers (recompute N
whole transformer layers) is incompatible with TE-scoped graphs. MCore
calls this "full" granularity even though you're selecting how many
layers — the name refers to recomputing the full layer, not full model.
Any TE-scoped scope (attn, mlp, moe_router, etc.) will assert:
AssertionError: full recompute is only supported with full iteration CUDA graph.
This commonly hits FP8 configs that default to TE-scoped graphs (e.g.
LLAMA3_70B_SFT_CONFIG_H100_FP8_CS_V1 uses cuda_graph_impl= "transformer_engine", cuda_graph_scope="mlp"). Fix: use submodule
recompute (recompute_granularity="selective" + recompute_modules),
disable CUDA graphs, or switch to local + full_iteration. Enforced
in MCore transformer_config.py:2001-2005. See also
@skills/nemo-mbridge-perf-activation-recompute/SKILL.md.
Benchmark numbers are workload-specific: graph wins are usually real when host overhead is visible, but the exact gain depends on batch shape, PP depth, recompute, dispatcher backend, and whether the eager baseline was already optimized.
A successful capture is not a speedup guarantee: On 2026-05-18,
Qwen3 30B A3B H100 BF16 pretrain with the all-to-all dispatcher captured
TE-scoped attn,moe_router,moe_preprocess graphs successfully (48
graphable layers, about 6.9 s capture time on rank 0), but replay
iterations 5-8 averaged 42.00 s versus 41.36 s for eager. Treat
scoped graphs as a bring-up candidate and validate on the target stack.
uv run python -m pytest \
tests/unit_tests/training/test_config.py -k "cuda_graph" \
tests/unit_tests/training/test_comm_overlap.py -k "cuda_graph" \
tests/unit_tests/models/test_gpt_full_te_layer_autocast_spec.py -k "cuda_graph" -quv run python -m pytest \
tests/functional_tests/test_groups/recipes/test_llama_recipes_pretrain_cuda_graphs.py -qlocal and
transformer_engine implementations.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-cuda-graphs of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
Nemo Mbridge Perf Cuda Graphs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Cuda Graphs this skillNVIDIA/skills | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| Quark Env Preflightamd/Quark | 181 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Hyperpod Version Checkerawslabs/agent-plugins | 915 | — | ~910 | Automated safety check: Pass | Apache-2.0 | |
| Spark Environment Setupwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Oob Perf Analysisintel/torch-xpu-ops | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
amd/Quark
Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
wshobson/agents
Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).
intel/torch-xpu-ops
Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA.
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. Nemo Mbridge Perf Cuda Graphs is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.
Nemo Mbridge Perf Cuda Graphs fits situations like: tasks that involve Deep learning.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-cuda-graphs in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-cuda-graphs in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-cuda-graphs in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-cuda-graphs in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-cuda-graphs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-cuda-graphs, .gemini/skills/nemo-mbridge-perf-cuda-graphs, .github/skills/nemo-mbridge-perf-cuda-graphs and .opencode/skills/nemo-mbridge-perf-cuda-graphs in your project.
Going by SKILL.md and its folder, Nemo Mbridge Perf Cuda Graphs needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Cuda Graphs is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Cuda Graphs: Graphsignal (graphsignal/graphsignal, 257 stars), Quark Env Preflight (amd/Quark, 181 stars), Hyperpod Version Checker (awslabs/agent-plugins, 915 stars) and Spark Environment Setup (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.