Setup Workshop Nemoclaw
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-recipe-recommender --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-recipe-recommender .claude/skills/nemo-mbridge-recipe-recommender && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-recipe-recommender" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-recipe-recommender into .claude/skills/nemo-mbridge-recipe-recommender/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-recipe-recommender", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-recipe-recommenderType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-recipe-recommender --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-recipe-recommender .agents/skills/nemo-mbridge-recipe-recommender && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-recipe-recommender" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-recipe-recommender into .agents/skills/nemo-mbridge-recipe-recommender/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-recipe-recommender", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-recipe-recommender --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-recipe-recommender .cursor/skills/nemo-mbridge-recipe-recommender && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-recipe-recommender" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-recipe-recommender into .cursor/skills/nemo-mbridge-recipe-recommender/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-recipe-recommender", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-recipe-recommender--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-recipe-recommender --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-recipe-recommender .gemini/skills/nemo-mbridge-recipe-recommender && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-recipe-recommender" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-recipe-recommender into .gemini/skills/nemo-mbridge-recipe-recommender/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-recipe-recommender", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-recipe-recommenderInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-recipe-recommender .github/skills/nemo-mbridge-recipe-recommender && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-recipe-recommender" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-recipe-recommender into .github/skills/nemo-mbridge-recipe-recommender/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-recipe-recommender", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-recipe-recommender --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-recipe-recommender .opencode/skills/nemo-mbridge-recipe-recommender && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-recipe-recommender" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-recipe-recommender into .opencode/skills/nemo-mbridge-recipe-recommender/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-recipe-recommender", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-recipe-recommenderRecommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal.
Nemo Mbridge Recipe Recommender is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal. Use when selecting a starting recipe, comparing library and benchmark configs, resizing parallelism for a GPU allocation, or distinguishing convergence changes, semantics-preserving execution tuning, and benchmark-only shortcuts.
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/recipe-index.md`).
It sits in AI & LLM Engineering, covering Fine-tuning. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Recipe Recommender loads about 4.1k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 1,582 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,582 words, ~4,138 tokens.
.claude/skills/nemo-mbridge-recipe-recommender/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.This skill indexes every shipped recipe and helps users pick the right starting config, adjust parallelism, and avoid common pitfalls.
When recommending recipes, always include these distinctions before the long index details:
src/megatron/bridge/recipes/ are for functional
training and use scripts/training/run_recipe.py.src/megatron/bridge/perf_recipes/ are for
upper-bound throughput benchmarks. They own their canonical benchmark data
and settings and should not be presented as production training recipes.llama3_8b_pretrain_config
with mock data via --dataset mock.--dataset squad or --dataset tulu3; for pretrain and mock validation
recommendations, use --dataset mock. Do not pair the pretraining-only
mock preset with an SFT or PEFT mode.num_key_value_heads, keep TP within one node unless using
NVL72-class interconnect, enable SP when TP > 1, configure CP for long
context, DP is implicit, and reduce micro_batch_size first on OOM.Separate training semantics from their hardware mapping before recommending or tuning a recipe.
Convergence configuration includes the starting checkpoint and trainable parameters; dataset/revision/split/order/seeds; tokenizer, masking, truncation, and packing; sequence length; global batch and token budget; objective and loss coefficients; natural or forced MoE routing and token-dropping policy; optimizer, LR, schedule, warmup, betas, epsilon, weight decay, clipping, and dropout; arithmetic and optimizer-state precision; and PEFT adapter settings. Changing one of these creates a new convergence experiment.
Execution/performance configuration includes hardware count and topology; TP/PP/VP/CP/EP/ETP/DP/SP; recompute and offload; distributed optimizer/FSDP; communication overlap; fusions and attention backends; CUDA graphs and compilation; checkpoint I/O; and MoE transport through all-to-all, DeepEP, or HybridEP when the routing policy is unchanged. These settings should preserve the objective and effective updates, although floating-point reduction order can produce small numerical drift that still needs validation.
Treat micro batch size and gradient accumulation as execution fingerprints. Tune them only with fixed global batch size, global batch membership/order, normalization, optimizer boundaries, and token budget, and validate fresh loss sentinels for each layout. Packing, precision, forced MoE load balancing, token dropping/capacity, and router/auxiliary loss changes are never performance-only knobs.
Treat mock data, forced balancing, disabled correctness checks, and timing-only
schedules as benchmark-only shortcuts. They may be appropriate in
perf_recipes, but their losses and checkpoints are not convergence evidence.
For comparable model-verification recipes, choose a cohort-wide convergence contract before tuning performance. Keep the same bounded data selection, preprocessing, sequence length, global batch, optimizer/schedule, precision, seeds, routing policy, optimizer-step horizon, and processed-token checkpoints where the architectures permit. Record any necessary model-specific deviation and do not present that result as apples-to-apples convergence evidence. Absolute losses from different architectures or tokenizers are not directly rankable; compare stability and trend at equal token counts.
When a recipe's batch disagrees with the chosen convergence contract, modify and validate the library recipe separately. A declared bounded-verification protocol may explicitly apply the same LR, schedule, sequence, and data overrides across a cohort, but do not make one-off convergence changes merely to improve throughput. Conversely, first try TP/PP/CP/EP, recompute/offload, dispatcher transport, overlap, fusion, and CUDA graphs when optimizing fit or throughput.
# Pretrain with mock data
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
--recipe <recipe_function_name> \
--dataset mock
# SFT with SQuAD
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
--recipe <recipe_function_name> \
--dataset squad
# Override any field via CLI
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
--recipe llama3_8b_pretrain_config \
--dataset mock \
'model.tensor_model_parallel_size=2' \
'train.global_batch_size=64'./scripts/training/train.sh \
--nodes 2 --gpus-per-node 8 \
--account ACCOUNT --partition PARTITION --container-image IMAGE \
--recipe qwen3_30b_a3b_pretrain_16gpu_h100_bf16_config \
--mode pretrainThe total GPU allocation must match the count encoded in the recipe name. The
user selects the node shape, and the selected partition must provide the
requested hardware. The launcher does not inject benchmark offline defaults or
cluster-specific launch policy. Use --env NAME for exported offline or NCCL
fabric settings and repeated --srun-arg=ARG options for srun.
Configure CPU/NUMA wrappers and Slurm segment sizing through the target cluster
integration, or use scripts/performance/setup_experiment.py when its compatibility
policies are required. The unified
launcher supports exact exported text pretraining, text SFT/PEFT, Qwen-VL
pretraining, and Wan pretraining recipes and infers their forward step. Text
SFT/PEFT text benchmark recipes retain the flat runner's mock-data default;
Qwen-VL and Wan retain their model-specific datasets. Exported benchmark PEFT
recipes are fixed LoRA configs; use a configurable library recipe for DoRA.
Trailing KEY=VALUE overrides are accepted, but an overridden benchmark
recipe no longer represents its canonical benchmark configuration. Use
scripts/performance/setup_experiment.py for selector-based invocation,
dataset replacement, topology resizing, and specialized benchmark controls.
See the Benchmark Recipe Index for important caveats before using these for anything beyond throughput benchmarking.
Benchmark recipes use the same Python function format as library recipes, but live in a dedicated namespace for throughput benchmarking:
src/megatron/bridge/perf_recipes/<family>/<hardware>/<model>.pyllama3_8b_pretrain_8gpu_h100_bf16_config())scripts/performance/utils/utils.py derives compatibility WorkloadBaseConfig views from the flat recipe itself_benchmark_common() (50 iters, timing, TE RNG), _perf_precision() (bf16 / fp8_cs / fp8_mx / nvfp4)Why Python, not YAML? Previous YAML-based approaches had problems: recipe logic was split across multiple indirection layers, configs were not self-contained, and the two-level pipeline made maintenance and debugging difficult. Python functions are explicit, greppable, and composable.
The training launcher discovers library and benchmark recipes from the complete exported function name. Five legacy duplicate names select the benchmark definition; use the corresponding generic alias for those functional workloads. New recipe names should be unique across both packages.
The full per-family recipe tables — every shipped library recipe
(src/megatron/bridge/recipes/) and benchmark recipe
(src/megatron/bridge/perf_recipes/), with parallelism degrees, minimum GPU
counts, and hardware coverage — are kept in a dedicated reference file so this
skill stays concise:
→ See references/recipe-index.md — Library
Recipe Index (Llama, Qwen2/2.5/3, Qwen3-MoE, Qwen3-Next, DeepSeek, GLM-4.5,
Gemma, Nemotron, VLM, Diffusion) and Benchmark Recipe Index (per-hardware
throughput configs).
Load that file to pull an exact recipe function name or its default parallelism; the guidance below tells you which entry to look up.
User wants to train a model
│
├─ Know the model name?
│ ├─ Yes → Look up in references/recipe-index.md
│ │ ├─ Has a recipe for their size + mode? → Use it directly
│ │ └─ No exact match? → Use closest size, adjust parallelism
│ └─ No → Ask for model name, size, and HF model ID
│
├─ What's the training goal?
│ ├─ Pretrain → Use *_pretrain_config
│ ├─ SFT (full fine-tune) → Use *_sft_config
│ └─ PEFT (LoRA/DoRA) → Use *_peft_config (lowest GPU requirement)
│
├─ How many GPUs?
│ ├─ 1 GPU → Only PEFT recipes work (TP=1, PP=1)
│ ├─ 8 GPUs (1 node) → Most 8B–16B models, small MoE (EP=8)
│ ├─ 16–64 GPUs → 70B dense, medium MoE
│ └─ 128+ GPUs → 405B+, large MoE (DeepSeek V3, Kimi K2)
│
├─ Want throughput benchmarks?
│ ├─ Yes → Use benchmark recipes (src/megatron/bridge/perf_recipes/)
│ │ ├─ Exact exported recipe → scripts/training/train.sh --recipe <exact function name>
│ │ └─ Selector/specialized workflow → scripts/performance/setup_experiment.py
│ └─ No → Use library recipes (scripts/training/run_recipe.py)
│
└─ Long context?
├─ > 8K → Need CP (context parallelism), check *_16k / *_64k / *_128k variants
└─ ≤ 8K → Default recipes workWhen the user's GPU count differs from the recipe default:
num_key_value_heads (GQA constraint). E.g. if
num_key_value_heads=8, valid TP = {1, 2, 4, 8}.cp_comm_type. For
GQA models, a2a+p2p hierarchical CP allows CP > num_kv_heads.PP × max(TP × CP, EP × ETP). Dense DP is
world_size / (TP × PP × CP) and expert EDP is
world_size / (PP × EP × ETP); both quotients must be integral, and the
expert count must be divisible by EP.micro_batch_size. If OOM, reduce to 1.global_batch_size determines learning dynamics. Scale with DP:
GBS = micro_batch_size × DP × gradient_accumulation_steps.micro_batch_size=1 is typical at scale.| Pitfall | Symptom | Fix |
|---|---|---|
| TP > num_kv_heads | Crash: "TP must divide num_query_groups" | Reduce TP to a divisor of num_kv_heads |
| PP without VP | Poor throughput (large bubble) | Set virtual_pipeline_model_parallel_size |
| EP too low for large MoE | OOM on expert params | Increase EP; each expert lives on EP/num_experts ranks |
| CUDA graphs + packed sequences | Assert: "CUDA graph accepts only Tensor inputs" | Disable packing or use local full-iteration graphs |
| CUDA graphs + full recompute | Assert: "full recompute only with full iteration CUDA graph" | Disable recompute or switch to local impl |
use_te_rng_tracker not set | Assert on provider init when CUDA graphs enabled | Set cfg.model.use_te_rng_tracker = True and cfg.rng.te_rng_tracker = True |
| FSDP + TP > 1 on H100 | Possible comm bottleneck | Prefer FSDP with TP=1 or TP=2 on H100; FSDP shines on GB/B-series |
| Long context without CP | OOM on activations | Add CP=2/4/8; use *_16k, *_64k, or *_128k recipe variants |
MoE overlap_grad_reduce on H100 | May hurt throughput (False in many H100 presets) | Set overlap_grad_reduce=False for MoE on H100 |
| VLM SFT missing image data | Runs but produces garbage | Provide actual multimodal dataset or use mock VLM data |
| Qwen35-VL MoE FSDP | Tested on Blackwell only | May not work on H100; validate first |
# Scale Llama3 8B from 2 GPUs to 8 GPUs (increase DP)
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
--recipe llama3_8b_pretrain_config \
--dataset mock
# Run the native 4-GPU Qwen3-MoE 30B PEFT topology
uv run python -m torch.distributed.run --nproc_per_node=4 scripts/training/run_recipe.py \
--recipe qwen3_30b_a3b_peft_config \
--dataset tulu3
# Add long context to an existing recipe
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
--recipe llama3_8b_pretrain_config \
--dataset mock \
'model.seq_length=32768' \
'model.context_parallel_size=4'
# Enable CUDA graphs on any recipe
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
--recipe qwen3_30b_a3b_pretrain_config \
--dataset mock \
'model.cuda_graph_impl=transformer_engine' \
'model.cuda_graph_scope=[attn,moe_router,moe_preprocess]' \
'model.use_te_rng_tracker=True' \
'rng.te_rng_tracker=True'| I want to... | Start with | GPUs needed |
|---|---|---|
| Try Bridge for the first time | llama3_8b_pretrain_config + mock data | 2 |
| Fine-tune a 7-8B model | llama3_8b_sft_config or qwen3_8b_sft_config | 2–4 |
| LoRA on 1 GPU | llama3_8b_peft_config or qwen3_8b_peft_config | 1 |
| Pretrain a dense 70B | llama3_70b_pretrain_config | 32–64 |
| Train a small MoE | qwen3_30b_a3b_pretrain_config | 16 |
| Train a large MoE (235B+) | qwen3_235b_a22b_pretrain_config | 256–512 |
| Benchmark text-pretrain throughput | Benchmark recipe via train.sh --recipe <exact name> | Exact encoded count |
| Long-context training | llama3_8b_128k_pretrain_config or add CP override | 16+ |
| VLM fine-tuning | qwen3_vl_8b_sft_config or gemma3_vl_*_sft_config | 4–8 |
| Diffusion training | wan_1_3B_pretrain_config or flux_12b_pretrain_config | 8 |
| What | Path |
|---|---|
| Library recipes root | src/megatron/bridge/recipes/ |
Recipe __init__.py (all exports) | src/megatron/bridge/recipes/__init__.py |
| Common recipe helpers | src/megatron/bridge/recipes/common.py |
| Training entry point | scripts/training/run_recipe.py |
| Training Slurm launcher | scripts/training/train.sh |
| Benchmark recipes root | src/megatron/bridge/perf_recipes/ |
| Benchmark compatibility launcher | scripts/performance/setup_experiment.py |
| Benchmark recipe helpers | scripts/performance/utils/utils.py |
| Benchmark overrides | scripts/performance/utils/overrides.py |
Last signature refresh: 2026-08-03.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/nemo-mbridge-recipe-recommender of NVIDIA/skills.
Open the folder on GitHubat commit dfdd080
Nemo Mbridge Recipe Recommender next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Recipe Recommender this skillNVIDIA/skills | 3.5k | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | |
| Nemotron Super3NVIDIA-NeMo/Nemotron | 2.1k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Nemotron 3 Ultra Text2sql LoraNVIDIA-NeMo/Nemotron | 2.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | |
| Nemotron UltraNVIDIA-NeMo/Nemotron | 2.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.7k | Automated safety check: Pass | MIT |
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
NVIDIA-NeMo/Nemotron
Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.
NVIDIA-NeMo/Nemotron
Run the Nemotron-3 Ultra Text2SQL LoRA fine-tuning tutorial (NeMo Megatron-Bridge) end-to-end for the user on their SLURM cluster: data prep, distributed checkpoint conversion, and packed LoRA…
NVIDIA-NeMo/Nemotron
Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.
Orchestra-Research/AI-Research-SKILLs
Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.
NVIDIA-NeMo/Nemotron
Add a new step under src/nemotron/steps/<category/<stepid/ — manifest (step.toml), runner glue, configs, and per-step README.md.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal. Nemo Mbridge Recipe Recommender is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal.
Nemo Mbridge Recipe Recommender fits situations like: selecting a starting recipe; comparing library and benchmark configs; resizing parallelism for a GPU allocation; distinguishing convergence changes.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-recipe-recommender in NVIDIA/skills) into .claude/skills/nemo-mbridge-recipe-recommender in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a codex`. Or copy the skill folder (skills/nemo-mbridge-recipe-recommender in NVIDIA/skills) into .agents/skills/nemo-mbridge-recipe-recommender in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-recipe-recommender, .gemini/skills/nemo-mbridge-recipe-recommender, .github/skills/nemo-mbridge-recipe-recommender and .opencode/skills/nemo-mbridge-recipe-recommender in your project.
Going by SKILL.md and its folder, Nemo Mbridge Recipe Recommender needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Recipe Recommender is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Nemo Mbridge Recipe Recommender: Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 146 stars), Nemotron Super3 (NVIDIA-NeMo/Nemotron, 2.1k stars), Nemotron 3 Ultra Text2sql Lora (NVIDIA-NeMo/Nemotron, 2.1k stars) and Nemotron Ultra (NVIDIA-NeMo/Nemotron, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.