Hyperpod Version Checker
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-run-on-slurm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mcore-run-on-slurm .claude/skills/mcore-run-on-slurm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "mcore-run-on-slurm" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-run-on-slurm into .claude/skills/mcore-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-run-on-slurm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-run-on-slurmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-run-on-slurm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/mcore-run-on-slurm .agents/skills/mcore-run-on-slurm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "mcore-run-on-slurm" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-run-on-slurm into .agents/skills/mcore-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-run-on-slurm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-run-on-slurm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/mcore-run-on-slurm .cursor/skills/mcore-run-on-slurm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "mcore-run-on-slurm" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-run-on-slurm into .cursor/skills/mcore-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-run-on-slurm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/Megatron-LM.git --path skills/mcore-run-on-slurm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-run-on-slurm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/mcore-run-on-slurm .gemini/skills/mcore-run-on-slurm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "mcore-run-on-slurm" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-run-on-slurm into .gemini/skills/mcore-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-run-on-slurm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/Megatron-LM mcore-run-on-slurmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/mcore-run-on-slurm .github/skills/mcore-run-on-slurm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "mcore-run-on-slurm" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-run-on-slurm into .github/skills/mcore-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-run-on-slurm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-run-on-slurm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/mcore-run-on-slurm .opencode/skills/mcore-run-on-slurm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "mcore-run-on-slurm" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-run-on-slurm into .opencode/skills/mcore-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-run-on-slurm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
mcore-run-on-slurmShows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
The skill answers SLURM setup questions with a short list of constants before the full script: submit from a shared worktree path every node can see, run one srun task per node, and start workers through uv run python -m torch.distributed.run rather than bare torchrun. MASTER_ADDR comes from scontrol show hostnames, and NNODES, GPUS_PER_NODE and WORLD_SIZE are derived from the allocation.
CUDA_DEVICE_MAX_CONNECTIONS depends on hardware and parallelism: it is 1 on pre-Blackwell Hopper or Ampere with tensor or context parallelism above 1 and no FSDP, unneeded on Blackwell or GB200, must not be 1 with Torch-FSDP2 or Megatron-FSDP, and is 32 for overlap_moe_expert_parallel_comm. Prerequisites are SLURM submission rights to a GPU partition, a checkout on a filesystem all nodes see, and uv with uv sync run beforehand. Container conventions, monitoring and per-rank failure diagnosis follow, though the excerpt was cut off before them.
Read from SKILL.md and the folder at commit eb50202. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Megatron-LM on SLURM loads about 1.8k tokens when it runs. Until then it costs about 145 tokens; SKILL.md has 714 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/Megatron-LM at commit eb50202, republished under its Apache-2.0 licence (© NVIDIA). 714 words, ~1,804 tokens.
.claude/skills/mcore-run-on-slurm/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.For text-only SLURM setup questions, answer with these constants before the full script:
cd there in the
script before launching training.srun task per node and launch workers with
uv run python -m torch.distributed.run, not bare torchrun.MASTER_ADDR from
scontrol show hostnames "$SLURM_JOB_NODELIST" | head -n1, set MASTER_PORT,
NNODES=${SLURM_NNODES}, GPUS_PER_NODE=<GPUS_PER_NODE>, and
WORLD_SIZE=$((NNODES * GPUS_PER_NODE)).--nnodes, --nproc-per-node, --node-rank, --master-addr, and
--master-port to torch.distributed.run.CUDA_DEVICE_MAX_CONNECTIONS: pre-Blackwell Hopper/Ampere with TP>1 or CP>1
and non-FSDP uses 1; Blackwell/GB200 does not need it; Torch-FSDP2 or
Megatron-FSDP must not use 1; overlap_moe_expert_parallel_comm uses 32.uv installed; run uv sync --extra training --extra dev (or --extra lts) on the worktree once before submission so the .venv is materialized and visible to every node.Save as run_megatron.slurm in the worktree:
#!/bin/bash
#SBATCH --job-name=megatron
#SBATCH --account=<SLURM_ACCOUNT>
#SBATCH --partition=<SLURM_PARTITION>
#SBATCH --nodes=<NODES>
#SBATCH --ntasks-per-node=1
#SBATCH --gpus-per-node=<GPUS_PER_NODE>
#SBATCH --time=<HH:MM:SS>
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
set -euo pipefail
cd <MEGATRON_WORKTREE>
export MASTER_ADDR=$(scontrol show hostnames "$SLURM_JOB_NODELIST" | head -n1)
export MASTER_PORT=${MASTER_PORT:-29500}
export NNODES=${SLURM_NNODES}
export GPUS_PER_NODE=<GPUS_PER_NODE>
export WORLD_SIZE=$((NNODES * GPUS_PER_NODE))
# Set CUDA_DEVICE_MAX_CONNECTIONS only when your configuration requires it
# (see the section below). Example for pre-Blackwell with TP>1 or CP>1
# (non-FSDP):
# export CUDA_DEVICE_MAX_CONNECTIONS=1
srun --ntasks=${NNODES} --ntasks-per-node=1 bash -c '
# NODE_RANK comes from SLURM_NODEID with one task per node.
NODE_RANK=${SLURM_NODEID}
uv run python -m torch.distributed.run \
--nnodes='"${NNODES}"' \
--nproc-per-node='"${GPUS_PER_NODE}"' \
--node-rank=${NODE_RANK} \
--master-addr='"${MASTER_ADDR}"' \
--master-port='"${MASTER_PORT}"' \
pretrain_gpt.py \
<MEGATRON_ARGS>
'Submit:
mkdir -p logs && JOB_ID=$(sbatch --parsable run_megatron.slurm)
echo "Submitted ${JOB_ID}"cd to it in the script. All nodes must reach the same path on a shared filesystem (NFS, Lustre, or similar) — node-local paths will not be visible to peer ranks.torchrun worker group across all nodes; do not start independent single-node jobs.--nproc-per-node should equal the number of visible GPUs per node.The right value depends on your hardware and parallelism mode. Do not export it unconditionally:
1. The relevant code path asserts on this — you will get an assertion error if it is not 1, not a silent deadlock.1. Leave the env var unset, or set it to a value greater than 1.overlap_moe_expert_parallel_comm enabled: set to 32.Set it explicitly in the sbatch script when your configuration calls for it.
Many sites run Megatron-LM inside a container (enroot/pyxis on some clusters, singularity on others). If you do, the uv-managed .venv must live on a path that is visible from inside the container, and the container image must provide the CUDA / NCCL / torch versions the repo expects (see docker/.ngc_version.dev and .ngc_version.lts). The skeleton above stays the same; wrap the srun invocation with your scheduler's container flags (--container-image=…, --container-mounts=…, etc.).
squeue -j "$JOB_ID" -o "%.10i %.8T %.10M %.6D %R"
sacct -j "$JOB_ID" --format=JobID,State,ExitCode,Elapsed
scancel "$JOB_ID"If your training script writes a result artifact (a JSON metrics file from rank 0, a final checkpoint, etc.), poll for the artifact rather than waiting only on squeue state. Useful output usually appears before SLURM marks the job complete, and polling on the artifact lets you cancel the job as soon as it lands instead of holding the allocation until the timeout.
Scan stderr from every rank, not just rank 0. The earliest non-NCCL Python traceback is usually the root cause; later NCCL timeouts on other ranks are downstream symptoms of the first crash.
Classify quickly:
WORLD_SIZE = TP × DP × CP × PP and head-count divisibility (num_attention_heads % TP == 0).uv sync, or stale PYTHONPATH. Confirm cd <MEGATRON_WORKTREE> before launch.MASTER_ADDR resolution, and command consistency across ranks.uv sync before the first submission. If the venv is missing, every job rebuilds it from inside srun, costing minutes per job.CUDA_DEVICE_MAX_CONNECTIONS=1 blindly. The right value depends on hardware and parallelism mode (see the dedicated section above). Setting it to 1 with FSDP causes a different problem; on Blackwell it has no effect; on pre-Blackwell with TP>1 or CP>1 (non-FSDP) the code asserts, it does not deadlock.torchrun instead of uv run python -m torch.distributed.run. Bare torchrun may dispatch through a python interpreter that does not see venv packages, depending on how the venv is set up.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in skills/mcore-run-on-slurm of NVIDIA/Megatron-LM.
Open the folder on GitHubat commit eb50202
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in NVIDIA/Megatron-LM, which our catalogue first saw on October 7, 2026.
Megatron-LM on SLURM next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Megatron-LM on SLURM this skillNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Hyperpod Version Checkerawslabs/agent-plugins | 915 | — | ~910 | Automated safety check: Pass | Apache-2.0 | |
| DGX Spark Training Gotchaswshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| OpenVLA-OFT Fine-TuningOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.7k | Automated safety check: Pass | MIT | |
| GPU OptimizerMathews-Tom/armory | 328 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Paddle Design CompilerPaddlePaddle/Paddle | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 |
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
wshobson/agents
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
Orchestra-Research/AI-Research-SKILLs
Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
PaddlePaddle/Paddle
A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
NVIDIA/Megatron-LM
Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
NVIDIA/Megatron-LM
Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
NVIDIA/Megatron-LM
Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document.
NVIDIA/Megatron-LM
Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.
Works with
Categories
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. run rather than bare torchrun. MASTER_ADDR comes from scontrol show hostnames, and NNODES, GPUS_PER_NODE and WORLD_SIZE are derived from the allocation.
Megatron-LM on SLURM fits situations like: writing an sbatch script for multi-node Megatron-LM training; choosing the right CUDA_DEVICE_MAX_CONNECTIONS value for a GPU generation; setting MASTER_ADDR and WORLD_SIZE from a SLURM allocation; diagnosing a failing rank in a distributed run.
Run `npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a claude-code`. Or copy the skill folder (skills/mcore-run-on-slurm in NVIDIA/Megatron-LM) into .claude/skills/mcore-run-on-slurm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a codex`. Or copy the skill folder (skills/mcore-run-on-slurm in NVIDIA/Megatron-LM) into .agents/skills/mcore-run-on-slurm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/Megatron-LM --skill mcore-run-on-slurm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mcore-run-on-slurm, .gemini/skills/mcore-run-on-slurm, .github/skills/mcore-run-on-slurm and .opencode/skills/mcore-run-on-slurm in your project.
Going by SKILL.md and its folder, Megatron-LM on SLURM needs the command-line tools its instructions call (uv). Our summary lists: A SLURM cluster login with access to a GPU partition; Megatron-LM checked out on a filesystem shared by all nodes; uv installed, with uv sync run before submission.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Megatron-LM on SLURM is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Megatron-LM on SLURM: Hyperpod Version Checker (awslabs/agent-plugins, 915 stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars), OpenVLA-OFT Fine-Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and GPU Optimizer (Mathews-Tom/armory, 328 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/Megatron-LM, which has 18,094 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/Megatron-LM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.