Official agent skill

Nemo Mbridge Recipe Recommender

by NVIDIA in NVIDIA/skills

Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Recipe Recommender

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-recipe-recommender --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-recipe-recommender .claude/skills/nemo-mbridge-recipe-recommender && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-recipe-recommender
GitHub stars
3.5k
Token cost
~4.1k tokens
SKILL.md length
1,582 words
Files
6 (incl. references)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal.

  • Works in 4 steps: Ask the user for: model name/size, GPU… → Look up the best-match recipe in the… → Recommend the recipe function name +… → …
  • Selecting a starting recipe
  • SKILL.md covers How to Use This Skill, First Answer Checklist, Configuration Layers and… and Entry Points, plus 6 more sections
  • Calls uv

What it does

Nemo Mbridge Recipe Recommender is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal. Use when selecting a starting recipe, comparing library and benchmark configs, resizing parallelism for a GPU allocation, or distinguishing convergence changes, semantics-preserving execution tuning, and benchmark-only shortcuts.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/recipe-index.md`).

It sits in AI & LLM Engineering, covering Fine-tuning. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Selecting a starting recipe
  • Comparing library and benchmark configs
  • Resizing parallelism for a GPU allocation
  • Distinguishing convergence changes

Example prompts

  • “/nemo-mbridge-recipe-recommender”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Ask the user for: model name/size, GPU count & type, training goal
  2. Look up the best-match recipe in the index below.
  3. Recommend the recipe function name + entry-point command.
  4. Provide adjustment advice (parallelism resizing, batch tuning, pitfalls).

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Recipe Recommender loads about 4.1k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 1,582 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,582 words, ~4,138 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-recipe-recommender/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-recipe-recommender
description
Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal. Use when selecting a starting recipe, comparing library and benchmark configs, resizing parallelism for a GPU allocation, or distinguishing convergence changes, semantics-preserving execution tuning, and benchmark-only shortcuts.
license
Apache-2.0

Auto Recipe — Recipe Index & Recommendation

This skill indexes every shipped recipe and helps users pick the right starting config, adjust parallelism, and avoid common pitfalls.

How to Use This Skill

  1. Ask the user for: model name/size, GPU count & type, training goal (pretrain / SFT / PEFT), and sequence length (if non-default).
  2. Look up the best-match recipe in the index below.
  3. Recommend the recipe function name + entry-point command.
  4. Provide adjustment advice (parallelism resizing, batch tuning, pitfalls).

First Answer Checklist

When recommending recipes, always include these distinctions before the long index details:

  1. Library recipes under src/megatron/bridge/recipes/ are for functional training and use scripts/training/run_recipe.py.
  2. Benchmark recipes under src/megatron/bridge/perf_recipes/ are for upper-bound throughput benchmarks. They own their canonical benchmark data and settings and should not be presented as production training recipes.
  3. For a first-time Bridge smoke test, recommend llama3_8b_pretrain_config with mock data via --dataset mock.
  4. For normal SFT recommendations, select a finetuning preset such as --dataset squad or --dataset tulu3; for pretrain and mock validation recommendations, use --dataset mock. Do not pair the pretraining-only mock preset with an SFT or PEFT mode.
  5. After the recipe and dataset, give the required resizing rules: TP must divide num_key_value_heads, keep TP within one node unless using NVL72-class interconnect, enable SP when TP > 1, configure CP for long context, DP is implicit, and reduce micro_batch_size first on OOM.
  6. State whether each proposed override changes the convergence contract or only the execution/performance mapping. Do not trade convergence semantics for throughput without calling it a new experiment.

Configuration Layers and Change Control

Separate training semantics from their hardware mapping before recommending or tuning a recipe.

Convergence configuration includes the starting checkpoint and trainable parameters; dataset/revision/split/order/seeds; tokenizer, masking, truncation, and packing; sequence length; global batch and token budget; objective and loss coefficients; natural or forced MoE routing and token-dropping policy; optimizer, LR, schedule, warmup, betas, epsilon, weight decay, clipping, and dropout; arithmetic and optimizer-state precision; and PEFT adapter settings. Changing one of these creates a new convergence experiment.

Execution/performance configuration includes hardware count and topology; TP/PP/VP/CP/EP/ETP/DP/SP; recompute and offload; distributed optimizer/FSDP; communication overlap; fusions and attention backends; CUDA graphs and compilation; checkpoint I/O; and MoE transport through all-to-all, DeepEP, or HybridEP when the routing policy is unchanged. These settings should preserve the objective and effective updates, although floating-point reduction order can produce small numerical drift that still needs validation.

Treat micro batch size and gradient accumulation as execution fingerprints. Tune them only with fixed global batch size, global batch membership/order, normalization, optimizer boundaries, and token budget, and validate fresh loss sentinels for each layout. Packing, precision, forced MoE load balancing, token dropping/capacity, and router/auxiliary loss changes are never performance-only knobs.

Treat mock data, forced balancing, disabled correctness checks, and timing-only schedules as benchmark-only shortcuts. They may be appropriate in perf_recipes, but their losses and checkpoints are not convergence evidence.

For comparable model-verification recipes, choose a cohort-wide convergence contract before tuning performance. Keep the same bounded data selection, preprocessing, sequence length, global batch, optimizer/schedule, precision, seeds, routing policy, optimizer-step horizon, and processed-token checkpoints where the architectures permit. Record any necessary model-specific deviation and do not present that result as apples-to-apples convergence evidence. Absolute losses from different architectures or tokenizers are not directly rankable; compare stability and trend at equal token counts.

When a recipe's batch disagrees with the chosen convergence contract, modify and validate the library recipe separately. A declared bounded-verification protocol may explicitly apply the same LR, schedule, sequence, and data overrides across a cohort, but do not make one-off convergence changes merely to improve throughput. Conversely, first try TP/PP/CP/EP, recompute/offload, dispatcher transport, overlap, fusion, and CUDA graphs when optimizing fit or throughput.


Entry Points

Library recipes (functional training)
bash
# Pretrain with mock data
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
    --recipe <recipe_function_name> \
    --dataset mock

# SFT with SQuAD
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
    --recipe <recipe_function_name> \
    --dataset squad

# Override any field via CLI
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
    --recipe llama3_8b_pretrain_config \
    --dataset mock \
    'model.tensor_model_parallel_size=2' \
    'train.global_batch_size=64'
Benchmark recipes (throughput benchmarks)
bash
./scripts/training/train.sh \
    --nodes 2 --gpus-per-node 8 \
    --account ACCOUNT --partition PARTITION --container-image IMAGE \
    --recipe qwen3_30b_a3b_pretrain_16gpu_h100_bf16_config \
    --mode pretrain

The total GPU allocation must match the count encoded in the recipe name. The user selects the node shape, and the selected partition must provide the requested hardware. The launcher does not inject benchmark offline defaults or cluster-specific launch policy. Use --env NAME for exported offline or NCCL fabric settings and repeated --srun-arg=ARG options for srun. Configure CPU/NUMA wrappers and Slurm segment sizing through the target cluster integration, or use scripts/performance/setup_experiment.py when its compatibility policies are required. The unified launcher supports exact exported text pretraining, text SFT/PEFT, Qwen-VL pretraining, and Wan pretraining recipes and infers their forward step. Text SFT/PEFT text benchmark recipes retain the flat runner's mock-data default; Qwen-VL and Wan retain their model-specific datasets. Exported benchmark PEFT recipes are fixed LoRA configs; use a configurable library recipe for DoRA. Trailing KEY=VALUE overrides are accepted, but an overridden benchmark recipe no longer represents its canonical benchmark configuration. Use scripts/performance/setup_experiment.py for selector-based invocation, dataset replacement, topology resizing, and specialized benchmark controls.

See the Benchmark Recipe Index for important caveats before using these for anything beyond throughput benchmarking.


Benchmark Recipe Layout

Benchmark recipes use the same Python function format as library recipes, but live in a dedicated namespace for throughput benchmarking:

  • Benchmark recipes live in src/megatron/bridge/perf_recipes/<family>/<hardware>/<model>.py
  • Each benchmark recipe is a self-contained Python function (e.g. llama3_8b_pretrain_8gpu_h100_bf16_config())
  • Recipe names encode model, task, GPU count, hardware, precision, and optional variant
  • scripts/performance/utils/utils.py derives compatibility WorkloadBaseConfig views from the flat recipe itself
  • Shared helpers: _benchmark_common() (50 iters, timing, TE RNG), _perf_precision() (bf16 / fp8_cs / fp8_mx / nvfp4)

Why Python, not YAML? Previous YAML-based approaches had problems: recipe logic was split across multiple indirection layers, configs were not self-contained, and the two-level pipeline made maintenance and debugging difficult. Python functions are explicit, greppable, and composable.

The training launcher discovers library and benchmark recipes from the complete exported function name. Five legacy duplicate names select the benchmark definition; use the corresponding generic alias for those functional workloads. New recipe names should be unique across both packages.


Show full SKILL.md (635 more words)Show less

Recipe Index (Library & Benchmark)

The full per-family recipe tables — every shipped library recipe (src/megatron/bridge/recipes/) and benchmark recipe (src/megatron/bridge/perf_recipes/), with parallelism degrees, minimum GPU counts, and hardware coverage — are kept in a dedicated reference file so this skill stays concise:

→ See references/recipe-index.md — Library Recipe Index (Llama, Qwen2/2.5/3, Qwen3-MoE, Qwen3-Next, DeepSeek, GLM-4.5, Gemma, Nemotron, VLM, Diffusion) and Benchmark Recipe Index (per-hardware throughput configs).

Load that file to pull an exact recipe function name or its default parallelism; the guidance below tells you which entry to look up.


Recommendation Decision Tree

text
User wants to train a model
│
├─ Know the model name?
│   ├─ Yes → Look up in references/recipe-index.md
│   │   ├─ Has a recipe for their size + mode? → Use it directly
│   │   └─ No exact match? → Use closest size, adjust parallelism
│   └─ No → Ask for model name, size, and HF model ID
│
├─ What's the training goal?
│   ├─ Pretrain → Use *_pretrain_config
│   ├─ SFT (full fine-tune) → Use *_sft_config
│   └─ PEFT (LoRA/DoRA) → Use *_peft_config (lowest GPU requirement)
│
├─ How many GPUs?
│   ├─ 1 GPU → Only PEFT recipes work (TP=1, PP=1)
│   ├─ 8 GPUs (1 node) → Most 8B–16B models, small MoE (EP=8)
│   ├─ 16–64 GPUs → 70B dense, medium MoE
│   └─ 128+ GPUs → 405B+, large MoE (DeepSeek V3, Kimi K2)
│
├─ Want throughput benchmarks?
│   ├─ Yes → Use benchmark recipes (src/megatron/bridge/perf_recipes/)
│   │   ├─ Exact exported recipe → scripts/training/train.sh --recipe <exact function name>
│   │   └─ Selector/specialized workflow → scripts/performance/setup_experiment.py
│   └─ No → Use library recipes (scripts/training/run_recipe.py)
│
└─ Long context?
    ├─ > 8K → Need CP (context parallelism), check *_16k / *_64k / *_128k variants
    └─ ≤ 8K → Default recipes work

Adjustment Advice (When Recommending)

Parallelism Resizing Rules

When the user's GPU count differs from the recipe default:

  1. TP must divide num_key_value_heads (GQA constraint). E.g. if num_key_value_heads=8, valid TP = {1, 2, 4, 8}.
  2. TP should stay within a single node (NVLink). TP > 8 requires inter-node NVLink (e.g., GB200 NVL72).
  3. PP adds pipeline bubbles. Minimize PP; only increase when TP alone can't fit the model. Use VP (virtual pipeline) to mitigate bubble overhead.
  4. EP doesn't reduce dense-layer memory. Only expert parameters shard with EP. Shared attention/embeddings are replicated. For "OOM with MoE", increase EP first, not TP.
  5. SP should be True whenever TP > 1. It eliminates redundant activation copies and is essentially free.
  6. CP requires all-to-all or ring attention. Check cp_comm_type. For GQA models, a2a+p2p hierarchical CP allows CP > num_kv_heads.
  7. Dense and expert meshes overlap. Do not multiply TP and EP together. The minimum MoE world size is PP × max(TP × CP, EP × ETP). Dense DP is world_size / (TP × PP × CP) and expert EDP is world_size / (PP × EP × ETP); both quotients must be integral, and the expert count must be divisible by EP.
Batch Size Tuning
  • Start with the recipe's micro_batch_size. If OOM, reduce to 1.
  • global_batch_size determines learning dynamics. Scale with DP: GBS = micro_batch_size × DP × gradient_accumulation_steps.
  • For MoE, micro_batch_size=1 is typical at scale.
Common Pitfalls to Warn About
PitfallSymptomFix
TP > num_kv_headsCrash: "TP must divide num_query_groups"Reduce TP to a divisor of num_kv_heads
PP without VPPoor throughput (large bubble)Set virtual_pipeline_model_parallel_size
EP too low for large MoEOOM on expert paramsIncrease EP; each expert lives on EP/num_experts ranks
CUDA graphs + packed sequencesAssert: "CUDA graph accepts only Tensor inputs"Disable packing or use local full-iteration graphs
CUDA graphs + full recomputeAssert: "full recompute only with full iteration CUDA graph"Disable recompute or switch to local impl
use_te_rng_tracker not setAssert on provider init when CUDA graphs enabledSet cfg.model.use_te_rng_tracker = True and cfg.rng.te_rng_tracker = True
FSDP + TP > 1 on H100Possible comm bottleneckPrefer FSDP with TP=1 or TP=2 on H100; FSDP shines on GB/B-series
Long context without CPOOM on activationsAdd CP=2/4/8; use *_16k, *_64k, or *_128k recipe variants
MoE overlap_grad_reduce on H100May hurt throughput (False in many H100 presets)Set overlap_grad_reduce=False for MoE on H100
VLM SFT missing image dataRuns but produces garbageProvide actual multimodal dataset or use mock VLM data
Qwen35-VL MoE FSDPTested on Blackwell onlyMay not work on H100; validate first
Recipe Override Examples
bash
# Scale Llama3 8B from 2 GPUs to 8 GPUs (increase DP)
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
    --recipe llama3_8b_pretrain_config \
    --dataset mock

# Run the native 4-GPU Qwen3-MoE 30B PEFT topology
uv run python -m torch.distributed.run --nproc_per_node=4 scripts/training/run_recipe.py \
    --recipe qwen3_30b_a3b_peft_config \
    --dataset tulu3

# Add long context to an existing recipe
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
    --recipe llama3_8b_pretrain_config \
    --dataset mock \
    'model.seq_length=32768' \
    'model.context_parallel_size=4'

# Enable CUDA graphs on any recipe
uv run python -m torch.distributed.run --nproc_per_node=8 scripts/training/run_recipe.py \
    --recipe qwen3_30b_a3b_pretrain_config \
    --dataset mock \
    'model.cuda_graph_impl=transformer_engine' \
    'model.cuda_graph_scope=[attn,moe_router,moe_preprocess]' \
    'model.use_te_rng_tracker=True' \
    'rng.te_rng_tracker=True'

Quick Reference: Which Recipe for My Situation?

I want to...Start withGPUs needed
Try Bridge for the first timellama3_8b_pretrain_config + mock data2
Fine-tune a 7-8B modelllama3_8b_sft_config or qwen3_8b_sft_config2–4
LoRA on 1 GPUllama3_8b_peft_config or qwen3_8b_peft_config1
Pretrain a dense 70Bllama3_70b_pretrain_config32–64
Train a small MoEqwen3_30b_a3b_pretrain_config16
Train a large MoE (235B+)qwen3_235b_a22b_pretrain_config256–512
Benchmark text-pretrain throughputBenchmark recipe via train.sh --recipe <exact name>Exact encoded count
Long-context trainingllama3_8b_128k_pretrain_config or add CP override16+
VLM fine-tuningqwen3_vl_8b_sft_config or gemma3_vl_*_sft_config4–8
Diffusion trainingwan_1_3B_pretrain_config or flux_12b_pretrain_config8

Code Anchors

WhatPath
Library recipes rootsrc/megatron/bridge/recipes/
Recipe __init__.py (all exports)src/megatron/bridge/recipes/__init__.py
Common recipe helperssrc/megatron/bridge/recipes/common.py
Training entry pointscripts/training/run_recipe.py
Training Slurm launcherscripts/training/train.sh
Benchmark recipes rootsrc/megatron/bridge/perf_recipes/
Benchmark compatibility launcherscripts/performance/setup_experiment.py
Benchmark recipe helpersscripts/performance/utils/utils.py
Benchmark overridesscripts/performance/utils/overrides.py

Last signature refresh: 2026-08-03.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/nemo-mbridge-recipe-recommender of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • references/recipe-index.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemo Mbridge Recipe Recommender next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Recipe Recommender compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Recipe Recommender this skillNVIDIA/skills3.5k—~4.1kAutomated safety check: PassApache-2.0
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0
Nemotron Super3NVIDIA-NeMo/Nemotron2.1k—~2.4kAutomated safety check: PassApache-2.0
Nemotron 3 Ultra Text2sql LoraNVIDIA-NeMo/Nemotron2.1k—~1.9kAutomated safety check: PassApache-2.0
Nemotron UltraNVIDIA-NeMo/Nemotron2.1k—~1.8kAutomated safety check: PassApache-2.0
OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT

Similar skills

  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Nemotron Super3

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes.

    2.1k GitHub stars~2.4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Nemotron 3 Ultra Text2sql Lora

    NVIDIA-NeMo/Nemotron

    Run the Nemotron-3 Ultra Text2SQL LoRA fine-tuning tutorial (NeMo Megatron-Bridge) end-to-end for the user on their SLURM cluster: data prep, distributed checkpoint conversion, and packed LoRA…

    2.1k GitHub stars~1.9k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Nemotron Ultra

    NVIDIA-NeMo/Nemotron

    Reference desk for NVIDIA Nemotron 3 Ultra (550B-A55B) — architecture, NVFP4 pretraining, SFT, MOPD (multi-teacher on-policy distillation), MTP boosting, quantization, inference.

    2.1k GitHub stars~1.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • OpenVLA-OFT Fine-Tuning

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

    13k GitHub stars~3.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Nemotron Add Step

    NVIDIA-NeMo/Nemotron

    Add a new step under src/nemotron/steps/<category/<stepid/ — manifest (step.toml), runner glue, configs, and per-step README.md.

    2.1k GitHub stars~1.7k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemo Mbridge Recipe Recommender

What does Nemo Mbridge Recipe Recommender do?

Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal. Nemo Mbridge Recipe Recommender is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Recommend and customize Megatron Bridge library and benchmark recipes for a user's model, GPU count, hardware, sequence length, and pretrain/SFT/PEFT goal.

When should I use Nemo Mbridge Recipe Recommender?

Nemo Mbridge Recipe Recommender fits situations like: selecting a starting recipe; comparing library and benchmark configs; resizing parallelism for a GPU allocation; distinguishing convergence changes.

How do I install Nemo Mbridge Recipe Recommender in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-recipe-recommender in NVIDIA/skills) into .claude/skills/nemo-mbridge-recipe-recommender in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Recipe Recommender in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a codex`. Or copy the skill folder (skills/nemo-mbridge-recipe-recommender in NVIDIA/skills) into .agents/skills/nemo-mbridge-recipe-recommender in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Recipe Recommender in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-recipe-recommender -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-recipe-recommender, .gemini/skills/nemo-mbridge-recipe-recommender, .github/skills/nemo-mbridge-recipe-recommender and .opencode/skills/nemo-mbridge-recipe-recommender in your project.

What does Nemo Mbridge Recipe Recommender need to run?

Going by SKILL.md and its folder, Nemo Mbridge Recipe Recommender needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Recipe Recommender access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Recipe Recommender safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Recipe Recommender use?

Nemo Mbridge Recipe Recommender is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Recipe Recommender use?

About 4.1k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Nemo Mbridge Recipe Recommender?

Skills that share tags, products or a category with Nemo Mbridge Recipe Recommender: Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 146 stars), Nemotron Super3 (NVIDIA-NeMo/Nemotron, 2.1k stars), Nemotron 3 Ultra Text2sql Lora (NVIDIA-NeMo/Nemotron, 2.1k stars) and Nemotron Ultra (NVIDIA-NeMo/Nemotron, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Recipe Recommender?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.