Official agent skill

Megatron-LM on SLURM

by NVIDIA in NVIDIA/Megatron-LM

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Megatron-LM on SLURM

skills CLI
$ npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/Megatron-LM mcore-run-on-slurm --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mcore-run-on-slurm .claude/skills/mcore-run-on-slurm && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mcore-run-on-slurm
GitHub stars
18k
Token cost
~1.8k tokens
SKILL.md length
714 words
Files
5
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

  • Writing an sbatch script for multi-node Megatron-LM training
  • SKILL.md covers Answer-First Constants, Prerequisites, Minimal sbatch script and Multi-node rules, plus 5 more sections
  • Calls uv
  • Choosing the right CUDA_DEVICE_MAX_CONNECTIONS value for a GPU generation

What it does

The skill answers SLURM setup questions with a short list of constants before the full script: submit from a shared worktree path every node can see, run one srun task per node, and start workers through uv run python -m torch.distributed.run rather than bare torchrun. MASTER_ADDR comes from scontrol show hostnames, and NNODES, GPUS_PER_NODE and WORLD_SIZE are derived from the allocation.

CUDA_DEVICE_MAX_CONNECTIONS depends on hardware and parallelism: it is 1 on pre-Blackwell Hopper or Ampere with tensor or context parallelism above 1 and no FSDP, unneeded on Blackwell or GB200, must not be 1 with Torch-FSDP2 or Megatron-FSDP, and is 32 for overlap_moe_expert_parallel_comm. Prerequisites are SLURM submission rights to a GPU partition, a checkout on a filesystem all nodes see, and uv with uv sync run beforehand. Container conventions, monitoring and per-rank failure diagnosis follow, though the excerpt was cut off before them.

When your agent uses it

  • Writing an sbatch script for multi-node Megatron-LM training
  • Choosing the right CUDA_DEVICE_MAX_CONNECTIONS value for a GPU generation
  • Setting MASTER_ADDR and WORLD_SIZE from a SLURM allocation
  • Diagnosing a failing rank in a distributed run

Example prompts

  • “Write an sbatch script that runs Megatron-LM across four nodes with eight GPUs each.”
  • “Tell me the right CUDA_DEVICE_MAX_CONNECTIONS for Hopper with tensor parallelism and no FSDP.”
  • “Our multi-node job hangs on startup, so check the torch.distributed.run environment variables.”

Requirements

  • A SLURM cluster login with access to a GPU partition
  • Megatron-LM checked out on a filesystem shared by all nodes
  • uv installed, with uv sync run before submission

What it can do on your machine

Read from SKILL.md and the folder at commit eb50202. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Megatron-LM on SLURM loads about 1.8k tokens when it runs. Until then it costs about 145 tokens; SKILL.md has 714 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/Megatron-LM at commit eb50202, republished under its Apache-2.0 licence (© NVIDIA). 714 words, ~1,804 tokens.

Download SKILL.mdSave it as .claude/skills/mcore-run-on-slurm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
mcore-run-on-slurm
description
How to launch distributed Megatron-LM training jobs on a SLURM cluster. Covers a minimal sbatch skeleton, environment-variable setup for torch.distributed.run, CUDA_DEVICE_MAX_CONNECTIONS rules across hardware and parallelism modes, container conventions, monitoring, and per-rank failure diagnosis. Use when submitting a SLURM job; writing or debugging an sbatch script; configuring multi-node distributed training; setting MASTER_ADDR / MASTER_PORT / WORLD_SIZE; diagnosing a SLURM job failure; 'how do I run on the cluster', 'sbatch', 'multi-node training'.
license
Apache-2.0
metadata.author
Oliver Koenig <okoenig@nvidia.com>

Run Megatron-LM on SLURM

Answer-First Constants

For text-only SLURM setup questions, answer with these constants before the full script:

  • Submit from a shared worktree path visible to every node; cd there in the script before launching training.
  • Use one srun task per node and launch workers with uv run python -m torch.distributed.run, not bare torchrun.
  • Set MASTER_ADDR from scontrol show hostnames "$SLURM_JOB_NODELIST" | head -n1, set MASTER_PORT, NNODES=${SLURM_NNODES}, GPUS_PER_NODE=<GPUS_PER_NODE>, and WORLD_SIZE=$((NNODES * GPUS_PER_NODE)).
  • Pass --nnodes, --nproc-per-node, --node-rank, --master-addr, and --master-port to torch.distributed.run.
  • CUDA_DEVICE_MAX_CONNECTIONS: pre-Blackwell Hopper/Ampere with TP>1 or CP>1 and non-FSDP uses 1; Blackwell/GB200 does not need it; Torch-FSDP2 or Megatron-FSDP must not use 1; overlap_moe_expert_parallel_comm uses 32.

Prerequisites

  • A SLURM cluster login with submission rights to a GPU partition.
  • Megatron-LM checked out on a filesystem visible to all nodes in the allocation (NFS, Lustre, or similar). All nodes must reach the same paths for code, data, checkpoints, and output.
  • uv installed; run uv sync --extra training --extra dev (or --extra lts) on the worktree once before submission so the .venv is materialized and visible to every node.

Minimal sbatch script

Save as run_megatron.slurm in the worktree:

bash
#!/bin/bash
#SBATCH --job-name=megatron
#SBATCH --account=<SLURM_ACCOUNT>
#SBATCH --partition=<SLURM_PARTITION>
#SBATCH --nodes=<NODES>
#SBATCH --ntasks-per-node=1
#SBATCH --gpus-per-node=<GPUS_PER_NODE>
#SBATCH --time=<HH:MM:SS>
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

set -euo pipefail
cd <MEGATRON_WORKTREE>

export MASTER_ADDR=$(scontrol show hostnames "$SLURM_JOB_NODELIST" | head -n1)
export MASTER_PORT=${MASTER_PORT:-29500}
export NNODES=${SLURM_NNODES}
export GPUS_PER_NODE=<GPUS_PER_NODE>
export WORLD_SIZE=$((NNODES * GPUS_PER_NODE))

# Set CUDA_DEVICE_MAX_CONNECTIONS only when your configuration requires it
# (see the section below). Example for pre-Blackwell with TP>1 or CP>1
# (non-FSDP):
#   export CUDA_DEVICE_MAX_CONNECTIONS=1

srun --ntasks=${NNODES} --ntasks-per-node=1 bash -c '
  # NODE_RANK comes from SLURM_NODEID with one task per node.
  NODE_RANK=${SLURM_NODEID}
  uv run python -m torch.distributed.run \
    --nnodes='"${NNODES}"' \
    --nproc-per-node='"${GPUS_PER_NODE}"' \
    --node-rank=${NODE_RANK} \
    --master-addr='"${MASTER_ADDR}"' \
    --master-port='"${MASTER_PORT}"' \
    pretrain_gpt.py \
      <MEGATRON_ARGS>
'

Submit:

bash
mkdir -p logs && JOB_ID=$(sbatch --parsable run_megatron.slurm)
echo "Submitted ${JOB_ID}"

Multi-node rules

  • Submit from the worktree you intend to run, or cd to it in the script. All nodes must reach the same path on a shared filesystem (NFS, Lustre, or similar) — node-local paths will not be visible to peer ranks.
  • Use one torchrun worker group across all nodes; do not start independent single-node jobs.
  • --nproc-per-node should equal the number of visible GPUs per node.
  • Write checkpoints, tensorboard data, and structured logs to shared storage.

CUDA_DEVICE_MAX_CONNECTIONS

The right value depends on your hardware and parallelism mode. Do not export it unconditionally:

  • Pre-Blackwell (Hopper, Ampere) with TP>1 or CP>1, non-FSDP: set to 1. The relevant code path asserts on this — you will get an assertion error if it is not 1, not a silent deadlock.
  • Blackwell: not required; setting it has no effect.
  • Torch-FSDP2 or Megatron-FSDP: must NOT be 1. Leave the env var unset, or set it to a value greater than 1.
  • overlap_moe_expert_parallel_comm enabled: set to 32.

Set it explicitly in the sbatch script when your configuration calls for it.

Containers

Many sites run Megatron-LM inside a container (enroot/pyxis on some clusters, singularity on others). If you do, the uv-managed .venv must live on a path that is visible from inside the container, and the container image must provide the CUDA / NCCL / torch versions the repo expects (see docker/.ngc_version.dev and .ngc_version.lts). The skeleton above stays the same; wrap the srun invocation with your scheduler's container flags (--container-image=…, --container-mounts=…, etc.).

Show full SKILL.md (288 more words)Show less

Monitor and collect

bash
squeue -j "$JOB_ID" -o "%.10i %.8T %.10M %.6D %R"
sacct -j "$JOB_ID" --format=JobID,State,ExitCode,Elapsed
scancel "$JOB_ID"

If your training script writes a result artifact (a JSON metrics file from rank 0, a final checkpoint, etc.), poll for the artifact rather than waiting only on squeue state. Useful output usually appears before SLURM marks the job complete, and polling on the artifact lets you cancel the job as soon as it lands instead of holding the allocation until the timeout.

Failure diagnosis

Scan stderr from every rank, not just rank 0. The earliest non-NCCL Python traceback is usually the root cause; later NCCL timeouts on other ranks are downstream symptoms of the first crash.

Classify quickly:

  • OOM: record rank, phase (forward / backward / optimizer), batch size, sequence length, parallelism (TP/DP/CP/PP), and peak memory before adjusting.
  • Shape / divisibility error: check WORLD_SIZE = TP × DP × CP × PP and head-count divisibility (num_attention_heads % TP == 0).
  • Import error: wrong worktree, missing uv sync, or stale PYTHONPATH. Confirm cd <MEGATRON_WORKTREE> before launch.
  • NCCL failure with no Python traceback: verify allocation, port reachability, MASTER_ADDR resolution, and command consistency across ranks.

Common pitfalls

  • Forgetting uv sync before the first submission. If the venv is missing, every job rebuilds it from inside srun, costing minutes per job.
  • Writing logs to a node-local path that disappears at job exit. Always write to the shared filesystem.
  • Setting CUDA_DEVICE_MAX_CONNECTIONS=1 blindly. The right value depends on hardware and parallelism mode (see the dedicated section above). Setting it to 1 with FSDP causes a different problem; on Blackwell it has no effect; on pre-Blackwell with TP>1 or CP>1 (non-FSDP) the code asserts, it does not deadlock.
  • Running bare torchrun instead of uv run python -m torch.distributed.run. Bare torchrun may dispatch through a python interpreter that does not see venv packages, depending on how the venv is set up.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/mcore-run-on-slurm of NVIDIA/Megatron-LM.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit eb50202

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in NVIDIA/Megatron-LM, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Megatron-LM on SLURM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Megatron-LM on SLURM compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Megatron-LM on SLURM this skillNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Hyperpod Version Checkerawslabs/agent-plugins915—~910Automated safety check: PassApache-2.0
DGX Spark Training Gotchaswshobson/agents40k—~2kAutomated safety check: PassMIT
OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT
GPU OptimizerMathews-Tom/armory328—~3.5kAutomated safety check: NotesMIT
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0

Similar skills

  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    915 GitHub stars~910 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub stars~2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • OpenVLA-OFT Fine-Tuning

    Orchestra-Research/AI-Research-SKILLs

    Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

    13k GitHub stars~3.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    328 GitHub stars~3.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/Megatron-LM

All 14 skills in this repo
  • Official

    Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.

    18k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Megatron-LM CI/CD Guide

    NVIDIA/Megatron-LM

    Official

    Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.

    18k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Official

    Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.

    18k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Official

    Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document.

    18k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Official

    Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.

    18k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed

Questions about Megatron-LM on SLURM

What does Megatron-LM on SLURM do?

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. run rather than bare torchrun. MASTER_ADDR comes from scontrol show hostnames, and NNODES, GPUS_PER_NODE and WORLD_SIZE are derived from the allocation.

When should I use Megatron-LM on SLURM?

Megatron-LM on SLURM fits situations like: writing an sbatch script for multi-node Megatron-LM training; choosing the right CUDA_DEVICE_MAX_CONNECTIONS value for a GPU generation; setting MASTER_ADDR and WORLD_SIZE from a SLURM allocation; diagnosing a failing rank in a distributed run.

How do I install Megatron-LM on SLURM in Claude Code?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a claude-code`. Or copy the skill folder (skills/mcore-run-on-slurm in NVIDIA/Megatron-LM) into .claude/skills/mcore-run-on-slurm in your project. Claude Code loads it when a task matches its description.

How do I install Megatron-LM on SLURM in Codex?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a codex`. Or copy the skill folder (skills/mcore-run-on-slurm in NVIDIA/Megatron-LM) into .agents/skills/mcore-run-on-slurm in your project. Codex loads it when a task matches its description.

Can I use Megatron-LM on SLURM in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mcore-run-on-slurm, .gemini/skills/mcore-run-on-slurm, .github/skills/mcore-run-on-slurm and .opencode/skills/mcore-run-on-slurm in your project.

What does Megatron-LM on SLURM need to run?

Going by SKILL.md and its folder, Megatron-LM on SLURM needs the command-line tools its instructions call (uv). Our summary lists: A SLURM cluster login with access to a GPU partition; Megatron-LM checked out on a filesystem shared by all nodes; uv installed, with uv sync run before submission.

Does Megatron-LM on SLURM access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Megatron-LM on SLURM safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Megatron-LM on SLURM use?

Megatron-LM on SLURM is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Megatron-LM on SLURM use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Megatron-LM on SLURM?

Skills that share tags, products or a category with Megatron-LM on SLURM: Hyperpod Version Checker (awslabs/agent-plugins, 915 stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars), OpenVLA-OFT Fine-Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and GPU Optimizer (Mathews-Tom/armory, 328 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Megatron-LM on SLURM?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/Megatron-LM, which has 18,094 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/Megatron-LM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.