Megatron-LM on SLURM
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
$ npx skills add wshobson/agents --skill spark-training-gotchas -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents spark-training-gotchas --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-training-gotchas .claude/skills/spark-training-gotchas && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spark-training-gotchas" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-training-gotchas into .claude/skills/spark-training-gotchas/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-training-gotchas", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-training-gotchasType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill spark-training-gotchas -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents spark-training-gotchas --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-training-gotchas .agents/skills/spark-training-gotchas && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spark-training-gotchas" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-training-gotchas into .agents/skills/spark-training-gotchas/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-training-gotchas", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill spark-training-gotchas -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents spark-training-gotchas --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-training-gotchas .cursor/skills/spark-training-gotchas && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spark-training-gotchas" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-training-gotchas into .cursor/skills/spark-training-gotchas/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-training-gotchas", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/dgx-spark-ops/skills/spark-training-gotchas--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill spark-training-gotchas -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents spark-training-gotchas --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-training-gotchas .gemini/skills/spark-training-gotchas && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spark-training-gotchas" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-training-gotchas into .gemini/skills/spark-training-gotchas/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-training-gotchas", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents spark-training-gotchasInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill spark-training-gotchas -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-training-gotchas .github/skills/spark-training-gotchas && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spark-training-gotchas" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-training-gotchas into .github/skills/spark-training-gotchas/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-training-gotchas", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill spark-training-gotchas -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents spark-training-gotchas --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-training-gotchas .opencode/skills/spark-training-gotchas && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spark-training-gotchas" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-training-gotchas into .opencode/skills/spark-training-gotchas/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-training-gotchas", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spark-training-gotchasPreflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.
The GB10 chip in DGX Spark, with Grace Blackwell, SM121, 128GB of unified memory and an aarch64 CPU, has ten recurring failure modes named G1 to G10 so tooling can check them by number. The skill is meant to be read before a long run, and applies when a run will not start, when it runs out of memory while `nvidia-smi` shows headroom, when throughput drops partway, before any multi-hour job, when linking two Sparks, or when choosing between FP8 and NVFP4.
A quick table pairs each symptom with its fix. Examples are using a cu130 wheel or container for CUDA ABI errors, skipping the pip build of flash-attn, dropping the page cache for OOMs despite headroom, expecting a sustained power cap near 100W, budgeting 180 to 192 GB/s of memory bandwidth, staying on FP8 unless `sm_121a` is available, using a container when an environment breaks, and using DDP or FSDP but never tensor parallelism across two Sparks. `assets/preflight.sh` and `references/gotcha-checks.md` support the checks.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
pippython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
DGX Spark Training Gotchas loads about 2k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 952 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 952 words, ~1,983 tokens.
.claude/skills/spark-training-gotchas/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.DGX Spark's GB10 chip (Grace Blackwell, SM121, 128GB unified memory, aarch64) has ten recurring failure modes across launch, memory, thermals, bandwidth, and precision. Each is named G1–G10 so it can be checked by number — the numbering is load-bearing for tooling that runs these checks. Read this before a long run, not after hour six.
nvidia-smi still shows headroom.| # | Symptom | Fix |
|---|---|---|
| G1 | undefined symbol / segfault | cu130 wheel or container |
| G2 | flash-attn wrong backend used | skip pip build; monkeypatch on NGC |
| G3 | OOM despite headroom | drop page cache |
| G4 | throughput drop / reboot | expect ~100W sustained cap |
| G5 | memory-bound step slow | budget 180–192 GB/s |
| G6 | cache evicted mid-run | one GPU server at a time |
| G7 | NVFP4 slower than FP8 | stay FP8 unless sm_121a |
| G8 | playbook fails outright | check upstream issues |
| G9 | env breaks after install | use a container |
| G10 | 2-Spark TP hangs | DDP/FSDP only, never TP |
ImportError: undefined symbol naming a CUDA
function, or a segfault on the first .cuda() call.libcudart.so.12; Spark
ships CUDA 13. pip never checks CUDA ABI, so it surfaces
only at import or first kernel launch.references/gotcha-checks.md G1 — the wheel's
CUDA build tag.download.pytorch.org/whl/cu130 or
use a matched container.pip install flash-attn still fails/hangs.
Unsloth may also silently train flash-attn over an
explicitly requested SDPA.attn_implementation="sdpa".references/gotcha-checks.md G2 — is flash-attn
already present and working.references/gotcha-checks.md G2.nvidia-smi still reports free memory under the 128GB cap
— or, on some setups, [N/A] outright instead of a number.references/gotcha-checks.md G3 — read free -g
and /proc/meminfo, not nvidia-smi.sync; echo 3 > /proc/sys/vm/drop_caches — needs root, a
between-run reset, not a mid-training step.references/gotcha-checks.md G4 — sample
nvidia-smi --query-gpu=temperature.gpu,power.draw.references/gotcha-checks.md G5 — observed step
time vs. the measured range, not spec.gpu-memory-utilization<=0.5.references/gotcha-checks.md
G6 — other GPU-resident
processes and whether
capped.cvt.e2m1x2 unless kernels target
sm_121a; NVFP4 runs ~32% slower without it.references/gotcha-checks.md G7 — capability
reports (12, 1); does the build target sm_121a?sm_121a.references/gotcha-checks.md G8 — the playbook
repo's recent issues.github.com/NVIDIA/dgx-spark-playbooks issues
before trusting a recipe for an expensive run.pip install, or two "identical"
environments behave differently.references/gotcha-checks.md
G9 — container or bare pip?spark-environment-setup
for tag guidance) or Unsloth's container. If bare pip is
unavoidable, follow the NVIDIA install order, including
--no-deps on Unsloth.references/gotcha-checks.md G10 — the
configured parallelism strategy.The cheapest checks to run before anything else:
python3 -c "import torch; print(torch.version.cuda)" # expect 13.x (G1); NGC builds have no +cu130 tag — that's not a failureimport torch; print(torch.cuda.get_device_capability()) # expect (12, 1) (G7){ [ -f /.dockerenv -o -f /run/.containerenv ] || grep -qE 'docker|containerd' /proc/1/cgroup; } 2>/dev/null && echo container || echo unknown # G9assets/preflight.sh runs G1, G3, G4, G7, G9 and produces one
output line per gotcha in a fixed format: G-number first, then
PASS/FAIL/WARN where automatable, SKIP when unavailable, or
INFO: for a raw reading (G3, G4). Full commands:
references/gotcha-checks.md. See also
spark-environment-setup for the environment assumed working.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references, assets) in plugins/dgx-spark-ops/skills/spark-training-gotchas of wshobson/agents.
Open the folder on GitHubat commit 46891e7
DGX Spark Training Gotchas next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| DGX Spark Training Gotchas this skillwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Megatron-LM on SLURMNVIDIA/Megatron-LM | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| GPU OptimizerMathews-Tom/armory | 329 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| Hyperpod Version Checkerawslabs/agent-plugins | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 |
NVIDIA/Megatron-LM
Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
Orchestra-Research/AI-Research-SKILLs
Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
Works with
Categories
Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. The GB10 chip in DGX Spark, with Grace Blackwell, SM121, 128GB of unified memory and an aarch64 CPU, has ten recurring failure modes named G1 to G10 so tooling can check them by number. The skill is meant to be read before a long run, and applies when a run will not start, when it runs out of memory while `nvidia-smi` shows headroom, when throughput drops partway, before any multi-hour job, when linking two Sparks, or when choosing between FP8 and NVFP4.
DGX Spark Training Gotchas fits situations like: starting a multi-hour training job on a DGX Spark; debugging an import error or segfault at launch on GB10; finding out why a run hit out-of-memory with free memory showing; choosing a parallelism strategy for two linked Sparks.
Run `npx skills add wshobson/agents --skill spark-training-gotchas -a claude-code`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-training-gotchas in wshobson/agents) into .claude/skills/spark-training-gotchas in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill spark-training-gotchas -a codex`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-training-gotchas in wshobson/agents) into .agents/skills/spark-training-gotchas in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill spark-training-gotchas -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-training-gotchas, .gemini/skills/spark-training-gotchas, .github/skills/spark-training-gotchas and .opencode/skills/spark-training-gotchas in your project.
Going by SKILL.md and its folder, DGX Spark Training Gotchas needs a shell for the scripts in its folder and the command-line tools its instructions call (pip and python3). Our summary lists: An NVIDIA DGX Spark (GB10) system; A shell to run assets/preflight.sh.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
DGX Spark Training Gotchas is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with DGX Spark Training Gotchas: Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), GPU Optimizer (Mathews-Tom/armory, 329 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and Hyperpod Version Checker (awslabs/agent-plugins, 916 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.