Official agent skill

Nemo Mbridge Multi Node Slurm

by NVIDIA in NVIDIA/skills

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemo Mbridge Multi Node Slurm

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-multi-node-slurm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-multi-node-slurm .claude/skills/nemo-mbridge-multi-node-slurm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-multi-node-slurm
GitHub stars
3.5k
Token cost
~3.9k tokens
SKILL.md length
1,443 words
Files
6 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.

  • Works in 8 steps: Add SBATCH Headers → Convert to Multi-Node → Wrap in TRAIN_CMD + two-phase srun → …
  • Tasks that involve GPU and accelerator computing
  • SKILL.md covers First Answer Checklist, Two Approaches: srun-native vs…, Cluster Environment and srun-native Approach (Preferred), plus 5 more sections
  • Calls uv, pip and python; needs HF_TOKEN and GH_TOKEN

What it does

Nemo Mbridge Multi Node Slurm is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/templates.md`).

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with NVIDIA AI Platform and Python. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve GPU and accelerator computing

Example prompts

  • “/nemo-mbridge-multi-node-slurm”

Requirements

  • Python 3
  • A credential in WANDB_API_KEY
  • A credential in YOUR_GITHUB_TOKEN

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Add SBATCH Headers
  2. Convert to Multi-Node
  3. Wrap in TRAIN_CMD + two-phase srun
  4. Launch (two-phase)
  5. (Optional) Add Loss Extraction Footer
  6. Allocate the node
  7. Launch container shell
  8. Set up environment inside container

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip
    • python
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • GH_TOKEN
    • WANDB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Multi Node Slurm loads about 3.9k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 1,443 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,443 words, ~3,940 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-multi-node-slurm/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-multi-node-slurm
description
Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.
license
Apache-2.0
when_to_use
Writing or converting Slurm sbatch scripts, scaling to multiple nodes, debugging NCCL/launch failures, or investigating a commit that caused multi-node…

Multi-Node Slurm

Convert single-node uv run python -m torch.distributed.run commands into multi-node Slurm sbatch scripts with Enroot container support, and debug common multi-node failures.

First Answer Checklist

When converting or debugging Bridge multi-node jobs, answer in this order:

  1. Prefer the srun-native launch shape for Bridge scripts that reach initialize.py: #SBATCH --ntasks-per-node=8 and a direct srun ... uv run python <script> ... launch. Do not wrap these jobs in python -m torch.distributed.run.
  2. State that Bridge derives RANK, WORLD_SIZE, LOCAL_RANK, MASTER_ADDR, and MASTER_PORT from SLURM variables during initialize.py distributed init.
  3. Require shared paths and matching container mounts for the repo, data, logs, HF_HOME, UV_CACHE_DIR, and NEMO_HOME.
  4. For NCCL timeout reports, do these first-log checks before speculating:
    • grep for real errors while filtering warning/frame noise
    • inspect Failures: to find the first failed rank and node
    • grep for ncclUniqueId, timeout, or crash on rank 0

Two Approaches: srun-native vs uv run torch.distributed

Approachntasks-per-nodeProcess spawningBest for
srun-native (preferred)8Slurm spawns 8 tasks/nodeConversion, inference, Bridge scripts
uv run torch.distributed (legacy)1uv run python -m torch.distributed.run spawns 8 procs/nodeMLM pretrain_gpt.py

Prefer srun-native — simpler, avoids shell escaping issues with TRAIN_CMD. Megatron Bridge auto-derives RANK, WORLD_SIZE, LOCAL_RANK, MASTER_ADDR, MASTER_PORT from SLURM env vars (SLURM_PROCID, SLURM_NTASKS, SLURM_LOCALID, SLURM_NODELIST) via common_utils.py helpers called during initialize.py distributed init, so you never need to set them manually.

Cluster Environment

Use a shared filesystem for the repository, data, logs, HF_HOME, UV_CACHE_DIR, and NEMO_HOME. NEMO_HOME must not use the container-local default (/root/.cache/nemo) for multi-node SFT/PEFT jobs, because packed-sequence data prepared on node 0 must be visible to the other nodes.

Keep credentials out of sbatch templates and logs. Provide HF_TOKEN, GH_TOKEN, and WANDB_API_KEY through the scheduler environment or a restricted secrets file, and never hardcode token values in the script body. For copy-paste environment and sbatch templates, read references/templates.md.

Log Directory
text
<SHARED_FS>/logs/<job_name>_<suffix>

srun-native Approach (Preferred)

Slurm spawns all processes directly. No torch.distributed.run, no TRAIN_CMD escaping.

SBATCH Headers
bash
#SBATCH --job-name=<model>-<task>
#SBATCH --nodes=<NNODES>
#SBATCH --ntasks-per-node=8          # Slurm spawns 8 tasks per node
#SBATCH --gpus-per-node=8
#SBATCH --time=00:30:00
#SBATCH --account=<YOUR_ACCOUNT>
#SBATCH --partition=batch
#SBATCH --output=<SHARED_FS>/logs/<job_name>_%j.log
#SBATCH --exclusive
Build and Launch

Use a two-phase srun pattern: first run a single-process uv sync to populate the shared cache, then launch the full multi-node job. The full copy-paste version lives in references/templates.md.

srun-native Key Points
  • Phase 1 runs uv sync once on a single node/process, building all wheels into the shared cache on Lustre
  • Phase 2's uv sync is a fast no-op (everything is cached) — safe to run on all ranks without sleep guards
  • initialize.py + common_utils.py auto-set RANK, WORLD_SIZE, LOCAL_RANK, MASTER_ADDR, MASTER_PORT from SLURM env vars
  • Env vars like HF_TOKEN, HF_HOME, UV_CACHE_DIR exported at sbatch level are inherited by srun tasks
  • Reference: examples/models/glm/glm_45v/slurm_sft.sh, examples/models/minimax/minimax_m2/slurm_conversion.sh

uv run torch.distributed Approach (Legacy)

Use when the script requires torch.distributed.run (e.g., MLM pretrain_gpt.py) or when Bridge's initialize.py is not in the call path.

1. Add SBATCH Headers
bash
#SBATCH --job-name=<model>-<framework>
#SBATCH --nodes=<NNODES>
#SBATCH --ntasks-per-node=1          # ALWAYS 1 — torchrun handles per-node spawning
#SBATCH --gpus-per-node=8
#SBATCH --time=00:30:00
#SBATCH --account=<YOUR_ACCOUNT>
#SBATCH --partition=batch
#SBATCH --output=<SHARED_FS>/logs/<job_name>_%j.log
#SBATCH --exclusive

Critical: --ntasks-per-node=1, NOT 8. uv run python -m torch.distributed.run --nproc_per_node=8 spawns 8 processes per node. Using ntasks-per-node=8 causes EADDRINUSE port collisions (8 tasks x 8 procs = 64 per node).

2. Convert to Multi-Node

Replace single-node:

bash
uv run python -m torch.distributed.run --nproc_per_node=8 \
  <script> <args>

With multi-node (inside TRAIN_CMD string):

bash
uv run python -m torch.distributed.run \
  --nproc_per_node=8 \
  --nnodes=\${SLURM_JOB_NUM_NODES} \
  --node_rank=\${SLURM_NODEID} \
  <script> <args>

MASTER_ADDR and MASTER_PORT are auto-derived from SLURM env vars by initialize.py / common_utils.py — no need to set them.

3. Wrap in TRAIN_CMD + two-phase srun

Use the same two-phase pattern: first a single-process srun to warm the uv cache, then the full run.

Set runtime variables inside the container, but do not inject token values into a long bash -c string. Export credentials through the scheduler or source a restricted secrets file before the job starts. Keep HF_HOME, UV_CACHE_DIR, and NEMO_HOME on shared storage.

4. Launch (two-phase)

Use the two-phase launch template in references/templates.md, keeping #SBATCH --ntasks-per-node=1 for this legacy approach.

bash
echo "======================================"
echo "Done. Losses:"
echo "======================================"
grep -E "iteration\s+" "$LOGDIR/<prefix>_${SLURM_JOB_ID}.log" | grep -iE "lm loss|reduced_train_loss" | head -25

Interactive GPU Allocation (salloc + srun)

For ad-hoc testing (inference, conversion debugging), always follow these 3 steps:

Step 1: Allocate the node
bash
salloc --account <YOUR_ACCOUNT> -N 1 \
  -J <YOUR_ACCOUNT>-debug \
  -p interactive --gpus-per-node=8 -t 240
Step 2: Launch container shell
bash
srun --mpi=pmix --no-kill \
  --container-image $CONTAINER_IMAGE \
  --container-mounts $CONTAINER_MOUNTS \
  --account <YOUR_ACCOUNT> -N 1 \
  -J <YOUR_ACCOUNT>-debug \
  --no-container-mount-home --gpus-per-node=8 \
  -p interactive --pty bash
Step 3: Set up environment inside container
bash
export GH_TOKEN=<YOUR_GITHUB_TOKEN>
wandb login <YOUR_WANDB_KEY>
export HF_TOKEN=<YOUR_HF_TOKEN>
export HF_HOME=<SHARED_FS>/HF_HOME
export UV_CACHE_DIR="<SHARED_FS>/uv_cache"
export NEMO_HOME="<SHARED_FS>/cache/nemo"
uv sync

Then run commands with uv run (uses the synced virtualenv):

bash
uv run python -m torch.distributed.run --nproc_per_node=8 \
  examples/conversion/hf_to_megatron_generate_text.py \
  --hf_model_path <org>/<model> --prompt "What is AI?" --max_new_tokens 50 --ep 8

Pitfalls with interactive allocation:

ErrorCauseFix
Cannot find GPU specificationMissing --gpus-per-nodeAlways include --gpus-per-node=8 in both salloc and srun
invalid partition specified: pool0Wrong partition nameUse interactive for interactive, batch for sbatch. Check: sinfo --summarize
Invalid account or account/partition combinationPartition not available for accountCheck combos: sacctmgr -nP show assoc where user=$USER format=account,partition
Unable to create step for job... Requested node configuration is not available-w <node> conflicts with allocationRemove -w flag — HF cache is on shared filesystem, accessible from any node
uv: command not found inside containerContainer doesn't have uv pre-installedUse a container with uv pre-installed, or pip install uv
No space left on device during uv or pipContainer's /root/.cache/ is fullRedirect: export UV_CACHE_DIR=<SHARED_FS>/uv_cache
ModuleNotFoundError: No module named 'megatron.core.activations'Container's pre-installed megatron-core conflicts with local 3rdparty/Megatron-LMInstall local: pip install -e 3rdparty/Megatron-LM --no-deps --no-build-isolation

Debugging Multi-Node Failures

Quick Diagnosis

Check the log for these patterns (in order):

bash
# 1. Find the actual error (filter noise)
grep -a 'Error\|OOM\|CUDA out of memory\|FAILED\|Killed' job.log \
  | grep -v 'UserWarning\|AllocatorConfig\|transformer_engine\|frame\|srun: error'

# 2. Check which rank crashed first
grep -a 'Failures:' -A 20 job.log | head -25

# 3. Check for NCCL timeout
grep -a 'ncclUniqueId\|timeout\|crash on rank 0' job.log | head -5
Debugging Checklist

When a multi-node job fails:

  1. Check exit code: 1 = Python error, 9 = OOM killed, 143 = SIGTERM (timeout or cascade)
  2. Find first failure: Which task/node crashed first? Others get SIGTERM (143) as cascade
  3. grep the actual error: Filter out UserWarnings, NCCL frame dumps
  4. Check rank 0 specifically: Most save/export errors happen on rank 0
  5. Verify EP sizing: For MoE models, ensure num_experts / EP fits in GPU memory with headroom
  6. Try interactive first: Use salloc -N 2 -p interactive to iterate faster than sbatch queue
Show full SKILL.md (557 more words)Show less
NCCL Timeout at dist.barrier() — "crash on rank 0"

Symptom: All ranks on node 2+ show:

text
[rank8] is setting up NCCL communicator and retrieving ncclUniqueId from [0]
... wait timeout after 600000ms
This may indicate a possible application crash on rank 0

Root causes (check in order):

CauseHow to verifyFix
save_artifacts hangs on rank 0Error is in save_hf_weights → dist.barrier()Increase timeout: init_process_group("nccl", timeout=timedelta(minutes=60))
ImportError in custom model codegrep ImportError job.logCatch ImportError in save_artifacts (see below)
Rank 0 OOM during exportgrep 'OutOfMemory' job.logIncrease EP or nodes
Network issue between nodesError only on cross-node ranksCheck sinfo, try different nodes

The save_artifacts problem: When trust_remote_code=True, rank 0 runs save_artifacts() (downloads tokenizer, config, custom modeling code) while all other ranks skip directly to dist.barrier(). If save_artifacts is slow or crashes, other ranks timeout.

Fix for ImportError in save_artifacts (hf_pretrained/base.py):

python
# Change:
except OSError:
    pass
# To:
except (OSError, ImportError):
    pass
OOM for MoE Models

Symptom: torch.OutOfMemoryError: CUDA out of memory during model loading or forward pass.

Key insight: TP does NOT reduce expert memory. Only EP splits experts across GPUs.

Sizing formula:

text
experts_per_gpu = num_experts / EP
expert_memory_gb ≈ experts_per_gpu * expert_params * 2 / 1e9  (bf16)
total_per_gpu ≈ expert_memory_gb + attention_memory_gb + kv_cache_gb

MiniMax-M2 example (256 experts, ~230GB fp8 → ~460GB bf16):

ConfigNodesGPUsExperts/GPUResult
TP=2, EP=41864OOM (too many experts)
TP=2, EP=821632Works for roundtrip (weight-only), OOM for inference
TP=1, EP=1621616Works for inference
TP=2, EP=328648Comfortable for training

Rules of thumb:

  • Roundtrip (weight-only): can use more experts per GPU (~60GB model params OK)
  • Inference (forward pass + KV cache): needs headroom (~40GB model params max)
  • Training (activations + optimizer): needs even more headroom (~30GB model params max)
ModuleNotFoundError: No module named 'megatron.core.tensor_parallel'

Cause: Container's pre-installed megatron-core conflicts with local 3rdparty/Megatron-LM.

Fix: Add uv sync before running:

bash
CMD="if [ \"\$SLURM_LOCALID\" -eq 0 ]; then uv sync; else sleep 10; fi && "
CMD="${CMD}uv run --no-sync python <script> <args>"
FP8 Weight Mismatch in Roundtrip

Symptom: Roundtrip completes but shows ❌ for all expert weights and raises ValueError: Weight mismatch detected.

Cause: Original HF weights are FP8, Megatron stores in BF16. Exported weights are BF16. Comparison against original FP8 exceeds atol=1e-1.

This is expected for FP8 models. The conversion is correct; the comparison tolerance is insufficient for the FP8→BF16 precision gap.

WORLD_SIZE Not Set with srun

Symptom: Script exits with "must be launched with torchrun".

Cause: Scripts check os.environ.get("WORLD_SIZE") which torchrun sets but srun doesn't.

Fix: Also check SLURM_NTASKS:

python
if os.environ.get("WORLD_SIZE") is None and os.environ.get("SLURM_NTASKS") is None:
    sys.exit(1)

Bridge's common_utils.py helpers (called by initialize.py) populate env vars from SLURM:

python
if "RANK" not in os.environ:
    os.environ["RANK"] = str(get_rank_safe())          # uses SLURM_PROCID
if "WORLD_SIZE" not in os.environ:
    os.environ["WORLD_SIZE"] = str(get_world_size_safe())  # uses SLURM_NTASKS
if "MASTER_ADDR" not in os.environ:
    os.environ["MASTER_ADDR"] = get_master_addr_safe()     # parses SLURM_NODELIST
if "MASTER_PORT" not in os.environ:
    os.environ["MASTER_PORT"] = str(get_master_port_safe()) # derives from SLURM_JOB_ID

Key Gotchas

  1. Two-phase srun for uv sync: Run a single-process srun first to warm the cache, then the full multi-node srun. The second uv sync is a fast no-op since everything is already cached on the shared filesystem.

  2. --no-container-mount-home is an srun flag, NOT an #SBATCH directive.

  3. Escaping inside TRAIN_CMD: Since TRAIN_CMD is a double-quoted string, escape inner $ for Slurm variables that must expand at runtime (not sbatch time):

    • \${SLURM_PROCID}, \${SLURM_JOB_NUM_NODES}, \${SLURM_NODEID}
    • Host-side variables like $GH_TOKEN, $LOGDIR, $WORKDIR expand at sbatch time — no escaping needed.
  4. Bridge rm -rf nemo_experiments: Add before training to avoid stale checkpoint auto-resume.

  5. MLM needs PYTHONPATH: For pretrain_gpt.py scripts, add inside TRAIN_CMD:

    bash
    PYTHONPATH=${WORKDIR}/3rdparty/Megatron-LM:\${PYTHONPATH:-} \
  6. Node count heuristic: Total GPUs = NNODES * 8. Must satisfy: TP * PP * EP * DP >= total_GPUs where DP = total_GPUs / (TP * PP * EP).

  7. NEMO_HOME on shared filesystem for multi-node SFT: The default nemo cache (/root/.cache/nemo) is container-local. Multi-node SFT with packed sequences prepares .npy files on one node that are invisible to others. Set export NEMO_HOME=<SHARED_FS>/cache/nemo so packed data is shared. Without this, ranks on other nodes fail with TypeError: 'NoneType' object is not an iterator.

Full Templates and Command Bodies

For copyable sbatch scaffolding and Bridge/MLM-specific TRAIN_CMD bodies, read references/templates.md.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in skills/nemo-mbridge-multi-node-slurm of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • references/templates.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Nemo Mbridge Multi Node Slurm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Multi Node Slurm compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Multi Node Slurm this skillNVIDIA/skills3.5k—~3.9kAutomated safety check: PassApache-2.0
TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT
Triton SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Hyperpod Version Checkerawslabs/agent-plugins9121 repos~910Automated safety check: PassApache-2.0
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Dstackdstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.0

Similar skills

  • TensorRT-LLM Inference

    Orchestra-Research/AI-Research-SKILLs

    Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

    13k GitHub starsUsed in 5 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Triton Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    912 GitHub starsUsed in 1 repo~910 tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Dstack

    dstackai/dstack

    dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

    2.3k GitHub stars~6.2k tokensUpdated today
    DevOps & CloudAuto-check: warnings
  • OpenVLA-OFT Fine-Tuning

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

    13k GitHub starsUsed in 1 repo~3.7k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Nemo Mbridge Multi Node Slurm

What does Nemo Mbridge Multi Node Slurm do?

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Nemo Mbridge Multi Node Slurm is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.

When should I use Nemo Mbridge Multi Node Slurm?

Nemo Mbridge Multi Node Slurm fits situations like: tasks that involve GPU and accelerator computing.

How do I install Nemo Mbridge Multi Node Slurm in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-multi-node-slurm in NVIDIA/skills) into .claude/skills/nemo-mbridge-multi-node-slurm in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Multi Node Slurm in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a codex`. Or copy the skill folder (skills/nemo-mbridge-multi-node-slurm in NVIDIA/skills) into .agents/skills/nemo-mbridge-multi-node-slurm in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Multi Node Slurm in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-multi-node-slurm, .gemini/skills/nemo-mbridge-multi-node-slurm, .github/skills/nemo-mbridge-multi-node-slurm and .opencode/skills/nemo-mbridge-multi-node-slurm in your project.

What does Nemo Mbridge Multi Node Slurm need to run?

Going by SKILL.md and its folder, Nemo Mbridge Multi Node Slurm needs the command-line tools its instructions call (uv, pip, python and bash) and credentials named HF_TOKEN, GH_TOKEN and WANDB_API_KEY. Our summary lists: Python 3; A credential in WANDB_API_KEY; A credential in YOUR_GITHUB_TOKEN.

Does Nemo Mbridge Multi Node Slurm access the network?

SKILL.md contains no URLs. Its commands use uv and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Multi Node Slurm safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Multi Node Slurm use?

Nemo Mbridge Multi Node Slurm is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Multi Node Slurm use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.

What are the alternatives to Nemo Mbridge Multi Node Slurm?

Skills that share tags, products or a category with Nemo Mbridge Multi Node Slurm: TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars), Triton Skill (slowlyC/agent-gpu-skills, 169 stars), Hyperpod Version Checker (awslabs/agent-plugins, 912 stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Multi Node Slurm?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.