TensorRT-LLM Inference
Orchestra-Research/AI-Research-SKILLs
Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.
Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-multi-node-slurm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-multi-node-slurm .claude/skills/nemo-mbridge-multi-node-slurm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-multi-node-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-multi-node-slurm into .claude/skills/nemo-mbridge-multi-node-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-multi-node-slurm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-multi-node-slurmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-multi-node-slurm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-multi-node-slurm .agents/skills/nemo-mbridge-multi-node-slurm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-multi-node-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-multi-node-slurm into .agents/skills/nemo-mbridge-multi-node-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-multi-node-slurm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-multi-node-slurm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-multi-node-slurm .cursor/skills/nemo-mbridge-multi-node-slurm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-multi-node-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-multi-node-slurm into .cursor/skills/nemo-mbridge-multi-node-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-multi-node-slurm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-multi-node-slurm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-multi-node-slurm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-multi-node-slurm .gemini/skills/nemo-mbridge-multi-node-slurm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-multi-node-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-multi-node-slurm into .gemini/skills/nemo-mbridge-multi-node-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-multi-node-slurm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-multi-node-slurmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-multi-node-slurm .github/skills/nemo-mbridge-multi-node-slurm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-multi-node-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-multi-node-slurm into .github/skills/nemo-mbridge-multi-node-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-multi-node-slurm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-multi-node-slurm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-multi-node-slurm .opencode/skills/nemo-mbridge-multi-node-slurm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-multi-node-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-multi-node-slurm into .opencode/skills/nemo-mbridge-multi-node-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-multi-node-slurm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-multi-node-slurmConvert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.
Nemo Mbridge Multi Node Slurm is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Covers srun-native vs uv run torch.distributed approaches, container setup, NCCL timeouts, OOM sizing for MoE models, and interactive allocation.
Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `BENCHMARK.md`, `evals/evals.json` and `references/templates.md`).
It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with NVIDIA AI Platform and Python. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvpippythonbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv and pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENGH_TOKENWANDB_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Multi Node Slurm loads about 3.9k tokens when it runs, and up to ~5.1k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 1,443 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,443 words, ~3,940 tokens.
.claude/skills/nemo-mbridge-multi-node-slurm/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Convert single-node uv run python -m torch.distributed.run commands into multi-node Slurm sbatch scripts with Enroot container support, and debug common multi-node failures.
When converting or debugging Bridge multi-node jobs, answer in this order:
initialize.py: #SBATCH --ntasks-per-node=8 and a direct srun ... uv run python <script> ... launch. Do not wrap these jobs in
python -m torch.distributed.run.RANK, WORLD_SIZE, LOCAL_RANK,
MASTER_ADDR, and MASTER_PORT from SLURM variables during
initialize.py distributed init.HF_HOME, UV_CACHE_DIR, and NEMO_HOME.Failures: to find the first failed rank and nodencclUniqueId, timeout, or crash on rank 0| Approach | ntasks-per-node | Process spawning | Best for |
|---|---|---|---|
| srun-native (preferred) | 8 | Slurm spawns 8 tasks/node | Conversion, inference, Bridge scripts |
| uv run torch.distributed (legacy) | 1 | uv run python -m torch.distributed.run spawns 8 procs/node | MLM pretrain_gpt.py |
Prefer srun-native — simpler, avoids shell escaping issues with TRAIN_CMD. Megatron Bridge auto-derives RANK, WORLD_SIZE, LOCAL_RANK, MASTER_ADDR, MASTER_PORT from SLURM env vars (SLURM_PROCID, SLURM_NTASKS, SLURM_LOCALID, SLURM_NODELIST) via common_utils.py helpers called during initialize.py distributed init, so you never need to set them manually.
Use a shared filesystem for the repository, data, logs, HF_HOME, UV_CACHE_DIR, and NEMO_HOME. NEMO_HOME must not use the container-local default (/root/.cache/nemo) for multi-node SFT/PEFT jobs, because packed-sequence data prepared on node 0 must be visible to the other nodes.
Keep credentials out of sbatch templates and logs. Provide HF_TOKEN, GH_TOKEN, and WANDB_API_KEY through the scheduler environment or a restricted secrets file, and never hardcode token values in the script body. For copy-paste environment and sbatch templates, read references/templates.md.
<SHARED_FS>/logs/<job_name>_<suffix>Slurm spawns all processes directly. No torch.distributed.run, no TRAIN_CMD escaping.
#SBATCH --job-name=<model>-<task>
#SBATCH --nodes=<NNODES>
#SBATCH --ntasks-per-node=8 # Slurm spawns 8 tasks per node
#SBATCH --gpus-per-node=8
#SBATCH --time=00:30:00
#SBATCH --account=<YOUR_ACCOUNT>
#SBATCH --partition=batch
#SBATCH --output=<SHARED_FS>/logs/<job_name>_%j.log
#SBATCH --exclusiveUse a two-phase srun pattern: first run a single-process uv sync to populate the shared cache, then launch the full multi-node job. The full copy-paste version lives in references/templates.md.
uv sync once on a single node/process, building all wheels into the shared cache on Lustreuv sync is a fast no-op (everything is cached) — safe to run on all ranks without sleep guardsinitialize.py + common_utils.py auto-set RANK, WORLD_SIZE, LOCAL_RANK, MASTER_ADDR, MASTER_PORT from SLURM env varsHF_TOKEN, HF_HOME, UV_CACHE_DIR exported at sbatch level are inherited by srun tasksexamples/models/glm/glm_45v/slurm_sft.sh, examples/models/minimax/minimax_m2/slurm_conversion.shUse when the script requires torch.distributed.run (e.g., MLM pretrain_gpt.py) or when Bridge's initialize.py is not in the call path.
#SBATCH --job-name=<model>-<framework>
#SBATCH --nodes=<NNODES>
#SBATCH --ntasks-per-node=1 # ALWAYS 1 — torchrun handles per-node spawning
#SBATCH --gpus-per-node=8
#SBATCH --time=00:30:00
#SBATCH --account=<YOUR_ACCOUNT>
#SBATCH --partition=batch
#SBATCH --output=<SHARED_FS>/logs/<job_name>_%j.log
#SBATCH --exclusiveCritical: --ntasks-per-node=1, NOT 8. uv run python -m torch.distributed.run --nproc_per_node=8 spawns 8 processes per node. Using ntasks-per-node=8 causes EADDRINUSE port collisions (8 tasks x 8 procs = 64 per node).
Replace single-node:
uv run python -m torch.distributed.run --nproc_per_node=8 \
<script> <args>With multi-node (inside TRAIN_CMD string):
uv run python -m torch.distributed.run \
--nproc_per_node=8 \
--nnodes=\${SLURM_JOB_NUM_NODES} \
--node_rank=\${SLURM_NODEID} \
<script> <args>MASTER_ADDR and MASTER_PORT are auto-derived from SLURM env vars by initialize.py / common_utils.py — no need to set them.
Use the same two-phase pattern: first a single-process srun to warm the uv cache, then the full run.
Set runtime variables inside the container, but do not inject token values into a long bash -c string. Export credentials through the scheduler or source a restricted secrets file before the job starts. Keep HF_HOME, UV_CACHE_DIR, and NEMO_HOME on shared storage.
Use the two-phase launch template in references/templates.md, keeping #SBATCH --ntasks-per-node=1 for this legacy approach.
echo "======================================"
echo "Done. Losses:"
echo "======================================"
grep -E "iteration\s+" "$LOGDIR/<prefix>_${SLURM_JOB_ID}.log" | grep -iE "lm loss|reduced_train_loss" | head -25salloc + srun)For ad-hoc testing (inference, conversion debugging), always follow these 3 steps:
salloc --account <YOUR_ACCOUNT> -N 1 \
-J <YOUR_ACCOUNT>-debug \
-p interactive --gpus-per-node=8 -t 240srun --mpi=pmix --no-kill \
--container-image $CONTAINER_IMAGE \
--container-mounts $CONTAINER_MOUNTS \
--account <YOUR_ACCOUNT> -N 1 \
-J <YOUR_ACCOUNT>-debug \
--no-container-mount-home --gpus-per-node=8 \
-p interactive --pty bashexport GH_TOKEN=<YOUR_GITHUB_TOKEN>
wandb login <YOUR_WANDB_KEY>
export HF_TOKEN=<YOUR_HF_TOKEN>
export HF_HOME=<SHARED_FS>/HF_HOME
export UV_CACHE_DIR="<SHARED_FS>/uv_cache"
export NEMO_HOME="<SHARED_FS>/cache/nemo"
uv syncThen run commands with uv run (uses the synced virtualenv):
uv run python -m torch.distributed.run --nproc_per_node=8 \
examples/conversion/hf_to_megatron_generate_text.py \
--hf_model_path <org>/<model> --prompt "What is AI?" --max_new_tokens 50 --ep 8Pitfalls with interactive allocation:
| Error | Cause | Fix |
|---|---|---|
Cannot find GPU specification | Missing --gpus-per-node | Always include --gpus-per-node=8 in both salloc and srun |
invalid partition specified: pool0 | Wrong partition name | Use interactive for interactive, batch for sbatch. Check: sinfo --summarize |
Invalid account or account/partition combination | Partition not available for account | Check combos: sacctmgr -nP show assoc where user=$USER format=account,partition |
Unable to create step for job... Requested node configuration is not available | -w <node> conflicts with allocation | Remove -w flag — HF cache is on shared filesystem, accessible from any node |
uv: command not found inside container | Container doesn't have uv pre-installed | Use a container with uv pre-installed, or pip install uv |
No space left on device during uv or pip | Container's /root/.cache/ is full | Redirect: export UV_CACHE_DIR=<SHARED_FS>/uv_cache |
ModuleNotFoundError: No module named 'megatron.core.activations' | Container's pre-installed megatron-core conflicts with local 3rdparty/Megatron-LM | Install local: pip install -e 3rdparty/Megatron-LM --no-deps --no-build-isolation |
Check the log for these patterns (in order):
# 1. Find the actual error (filter noise)
grep -a 'Error\|OOM\|CUDA out of memory\|FAILED\|Killed' job.log \
| grep -v 'UserWarning\|AllocatorConfig\|transformer_engine\|frame\|srun: error'
# 2. Check which rank crashed first
grep -a 'Failures:' -A 20 job.log | head -25
# 3. Check for NCCL timeout
grep -a 'ncclUniqueId\|timeout\|crash on rank 0' job.log | head -5When a multi-node job fails:
num_experts / EP fits in GPU memory with headroomsalloc -N 2 -p interactive to iterate faster than sbatch queuedist.barrier() — "crash on rank 0"Symptom: All ranks on node 2+ show:
[rank8] is setting up NCCL communicator and retrieving ncclUniqueId from [0]
... wait timeout after 600000ms
This may indicate a possible application crash on rank 0Root causes (check in order):
| Cause | How to verify | Fix |
|---|---|---|
save_artifacts hangs on rank 0 | Error is in save_hf_weights → dist.barrier() | Increase timeout: init_process_group("nccl", timeout=timedelta(minutes=60)) |
ImportError in custom model code | grep ImportError job.log | Catch ImportError in save_artifacts (see below) |
| Rank 0 OOM during export | grep 'OutOfMemory' job.log | Increase EP or nodes |
| Network issue between nodes | Error only on cross-node ranks | Check sinfo, try different nodes |
The save_artifacts problem: When trust_remote_code=True, rank 0 runs save_artifacts() (downloads tokenizer, config, custom modeling code) while all other ranks skip directly to dist.barrier(). If save_artifacts is slow or crashes, other ranks timeout.
Fix for ImportError in save_artifacts (hf_pretrained/base.py):
# Change:
except OSError:
pass
# To:
except (OSError, ImportError):
passSymptom: torch.OutOfMemoryError: CUDA out of memory during model loading or forward pass.
Key insight: TP does NOT reduce expert memory. Only EP splits experts across GPUs.
Sizing formula:
experts_per_gpu = num_experts / EP
expert_memory_gb ≈ experts_per_gpu * expert_params * 2 / 1e9 (bf16)
total_per_gpu ≈ expert_memory_gb + attention_memory_gb + kv_cache_gbMiniMax-M2 example (256 experts, ~230GB fp8 → ~460GB bf16):
| Config | Nodes | GPUs | Experts/GPU | Result |
|---|---|---|---|---|
| TP=2, EP=4 | 1 | 8 | 64 | OOM (too many experts) |
| TP=2, EP=8 | 2 | 16 | 32 | Works for roundtrip (weight-only), OOM for inference |
| TP=1, EP=16 | 2 | 16 | 16 | Works for inference |
| TP=2, EP=32 | 8 | 64 | 8 | Comfortable for training |
Rules of thumb:
ModuleNotFoundError: No module named 'megatron.core.tensor_parallel'Cause: Container's pre-installed megatron-core conflicts with local 3rdparty/Megatron-LM.
Fix: Add uv sync before running:
CMD="if [ \"\$SLURM_LOCALID\" -eq 0 ]; then uv sync; else sleep 10; fi && "
CMD="${CMD}uv run --no-sync python <script> <args>"Symptom: Roundtrip completes but shows ❌ for all expert weights and raises ValueError: Weight mismatch detected.
Cause: Original HF weights are FP8, Megatron stores in BF16. Exported weights are BF16. Comparison against original FP8 exceeds atol=1e-1.
This is expected for FP8 models. The conversion is correct; the comparison tolerance is insufficient for the FP8→BF16 precision gap.
WORLD_SIZE Not Set with srunSymptom: Script exits with "must be launched with torchrun".
Cause: Scripts check os.environ.get("WORLD_SIZE") which torchrun sets but srun doesn't.
Fix: Also check SLURM_NTASKS:
if os.environ.get("WORLD_SIZE") is None and os.environ.get("SLURM_NTASKS") is None:
sys.exit(1)Bridge's common_utils.py helpers (called by initialize.py) populate env vars from SLURM:
if "RANK" not in os.environ:
os.environ["RANK"] = str(get_rank_safe()) # uses SLURM_PROCID
if "WORLD_SIZE" not in os.environ:
os.environ["WORLD_SIZE"] = str(get_world_size_safe()) # uses SLURM_NTASKS
if "MASTER_ADDR" not in os.environ:
os.environ["MASTER_ADDR"] = get_master_addr_safe() # parses SLURM_NODELIST
if "MASTER_PORT" not in os.environ:
os.environ["MASTER_PORT"] = str(get_master_port_safe()) # derives from SLURM_JOB_IDTwo-phase srun for uv sync: Run a single-process srun first to warm the cache, then the full multi-node srun. The second uv sync is a fast no-op since everything is already cached on the shared filesystem.
--no-container-mount-home is an srun flag, NOT an #SBATCH directive.
Escaping inside TRAIN_CMD: Since TRAIN_CMD is a double-quoted string, escape inner $ for Slurm variables that must expand at runtime (not sbatch time):
\${SLURM_PROCID}, \${SLURM_JOB_NUM_NODES}, \${SLURM_NODEID}$GH_TOKEN, $LOGDIR, $WORKDIR expand at sbatch time — no escaping needed.Bridge rm -rf nemo_experiments: Add before training to avoid stale checkpoint auto-resume.
MLM needs PYTHONPATH: For pretrain_gpt.py scripts, add inside TRAIN_CMD:
PYTHONPATH=${WORKDIR}/3rdparty/Megatron-LM:\${PYTHONPATH:-} \Node count heuristic: Total GPUs = NNODES * 8. Must satisfy: TP * PP * EP * DP >= total_GPUs where DP = total_GPUs / (TP * PP * EP).
NEMO_HOME on shared filesystem for multi-node SFT: The default nemo cache (/root/.cache/nemo)
is container-local. Multi-node SFT with packed sequences prepares .npy files on one node
that are invisible to others. Set export NEMO_HOME=<SHARED_FS>/cache/nemo so packed data
is shared. Without this, ranks on other nodes fail with TypeError: 'NoneType' object is not an iterator.
For copyable sbatch scaffolding and Bridge/MLM-specific TRAIN_CMD bodies, read
references/templates.md.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/nemo-mbridge-multi-node-slurm of NVIDIA/skills.
Open the folder on GitHubat commit 0e0d506
Nemo Mbridge Multi Node Slurm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Multi Node Slurm this skillNVIDIA/skills | 3.5k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| TensorRT-LLM InferenceOrchestra-Research/AI-Research-SKILLs | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Triton SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Hyperpod Version Checkerawslabs/agent-plugins | 912 | 1 repos | ~910 | Automated safety check: Pass | Apache-2.0 | |
| Megatron-LM on SLURMNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Dstackdstackai/dstack | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 |
Orchestra-Research/AI-Research-SKILLs
Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.
slowlyC/agent-gpu-skills
Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
dstackai/dstack
dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.
Orchestra-Research/AI-Research-SKILLs
Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. Nemo Mbridge Multi Node Slurm is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.
Nemo Mbridge Multi Node Slurm fits situations like: tasks that involve GPU and accelerator computing.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-multi-node-slurm in NVIDIA/skills) into .claude/skills/nemo-mbridge-multi-node-slurm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a codex`. Or copy the skill folder (skills/nemo-mbridge-multi-node-slurm in NVIDIA/skills) into .agents/skills/nemo-mbridge-multi-node-slurm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-multi-node-slurm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-multi-node-slurm, .gemini/skills/nemo-mbridge-multi-node-slurm, .github/skills/nemo-mbridge-multi-node-slurm and .opencode/skills/nemo-mbridge-multi-node-slurm in your project.
Going by SKILL.md and its folder, Nemo Mbridge Multi Node Slurm needs the command-line tools its instructions call (uv, pip, python and bash) and credentials named HF_TOKEN, GH_TOKEN and WANDB_API_KEY. Our summary lists: Python 3; A credential in WANDB_API_KEY; A credential in YOUR_GITHUB_TOKEN.
SKILL.md contains no URLs. Its commands use uv and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Multi Node Slurm is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Nemo Mbridge Multi Node Slurm: TensorRT-LLM Inference (Orchestra-Research/AI-Research-SKILLs, 13k stars), Triton Skill (slowlyC/agent-gpu-skills, 169 stars), Hyperpod Version Checker (awslabs/agent-plugins, 912 stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.