Agent skill

DGX Spark Training Gotchas

by wshobson in wshobson/agents

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

MITAuto-check passedAI & LLM Engineering

Install DGX Spark Training Gotchas

skills CLI
$ npx skills add wshobson/agents --skill spark-training-gotchas -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents spark-training-gotchas --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-training-gotchas .claude/skills/spark-training-gotchas && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spark-training-gotchas
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
952 words
Files
3 (incl. references, assets)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

  • Starting a multi-hour training job on a DGX Spark
  • SKILL.md covers When to Use This Skill, Common Issues Quick Reference, The Ten Gotchas and Fast Triage
  • Runs Shell scripts from its folder; calls pip and python3
  • Debugging an import error or segfault at launch on GB10

What it does

The GB10 chip in DGX Spark, with Grace Blackwell, SM121, 128GB of unified memory and an aarch64 CPU, has ten recurring failure modes named G1 to G10 so tooling can check them by number. The skill is meant to be read before a long run, and applies when a run will not start, when it runs out of memory while `nvidia-smi` shows headroom, when throughput drops partway, before any multi-hour job, when linking two Sparks, or when choosing between FP8 and NVFP4.

A quick table pairs each symptom with its fix. Examples are using a cu130 wheel or container for CUDA ABI errors, skipping the pip build of flash-attn, dropping the page cache for OOMs despite headroom, expecting a sustained power cap near 100W, budgeting 180 to 192 GB/s of memory bandwidth, staying on FP8 unless `sm_121a` is available, using a container when an environment breaks, and using DDP or FSDP but never tensor parallelism across two Sparks. `assets/preflight.sh` and `references/gotcha-checks.md` support the checks.

When your agent uses it

  • Starting a multi-hour training job on a DGX Spark
  • Debugging an import error or segfault at launch on GB10
  • Finding out why a run hit out-of-memory with free memory showing
  • Choosing a parallelism strategy for two linked Sparks

Example prompts

  • “Run the Spark preflight before I start this overnight fine-tune.”
  • “My training run segfaults on the first cuda call on the Spark, so diagnose it.”
  • “Throughput dropped halfway through the run, so check the known gotchas.”
  • “Should I use FP8 or NVFP4 for this run on GB10?”

Requirements

  • An NVIDIA DGX Spark (GB10) system
  • A shell to run assets/preflight.sh

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

DGX Spark Training Gotchas loads about 2k tokens when it runs, and up to ~4.4k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 952 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 952 words, ~1,983 tokens.

Download SKILL.mdSave it as .claude/skills/spark-training-gotchas/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
spark-training-gotchas
description
Preflight and diagnose the ten known failure modes for ML training on NVIDIA DGX Spark. Use when a training run on DGX Spark fails to start, OOMs below the 128GB limit, slows down mid-run, or before any multi-hour training job on GB10.

Spark Training Gotchas

DGX Spark's GB10 chip (Grace Blackwell, SM121, 128GB unified memory, aarch64) has ten recurring failure modes across launch, memory, thermals, bandwidth, and precision. Each is named G1–G10 so it can be checked by number — the numbering is load-bearing for tooling that runs these checks. Read this before a long run, not after hour six.

When to Use This Skill

  • A training run fails to start, with an import error or a segfault that doesn't point at the real cause.
  • A run OOMs while nvidia-smi still shows headroom.
  • Throughput degrades partway through a run that started fine.
  • Before any multi-hour or multi-epoch job on GB10.
  • Wiring two Sparks together, before picking a parallelism strategy.
  • Choosing between FP8 and NVFP4 for a Spark-hosted run.

Common Issues Quick Reference

#SymptomFix
G1undefined symbol / segfaultcu130 wheel or container
G2flash-attn wrong backend usedskip pip build; monkeypatch on NGC
G3OOM despite headroomdrop page cache
G4throughput drop / rebootexpect ~100W sustained cap
G5memory-bound step slowbudget 180–192 GB/s
G6cache evicted mid-runone GPU server at a time
G7NVFP4 slower than FP8stay FP8 unless sm_121a
G8playbook fails outrightcheck upstream issues
G9env breaks after installuse a container
G102-Spark TP hangsDDP/FSDP only, never TP

The Ten Gotchas

G1: CUDA 12/13 ABI Mismatch
  • SYMPTOM: ImportError: undefined symbol naming a CUDA function, or a segfault on the first .cuda() call.
  • CAUSE: most PyPI wheels link libcudart.so.12; Spark ships CUDA 13. pip never checks CUDA ABI, so it surfaces only at import or first kernel launch.
  • CHECK: references/gotcha-checks.md G1 — the wheel's CUDA build tag.
  • FIX: reinstall from download.pytorch.org/whl/cu130 or use a matched container.
G2: flash-attn — Skip the pip Build, Watch Unsloth's Auto-Detect
  • SYMPTOM: pip install flash-attn still fails/hangs. Unsloth may also silently train flash-attn over an explicitly requested SDPA.
  • CAUSE: no aarch64/sm_121 wheel for bare pip — but NGC containers ship a working SM121 flash-attn, and Unsloth auto-prefers it, dropping attn_implementation="sdpa".
  • CHECK: references/gotcha-checks.md G2 — is flash-attn already present and working.
  • FIX: bare pip — skip flash-attn, use SDPA (unchanged). On NGC — the only reliable override is the monkeypatch in references/gotcha-checks.md G2.
G3: UMA OOM Below 128GB
  • SYMPTOM: OOM during model load/training while nvidia-smi still reports free memory under the 128GB cap — or, on some setups, [N/A] outright instead of a number.
  • CAUSE: mmap and the CUDA allocator double-count pages during safetensors load; QLoRA can OOM earlier than bf16 since dequantization adds transient allocs.
  • CHECK: references/gotcha-checks.md G3 — read free -g and /proc/meminfo, not nvidia-smi.
  • FIX: drop the page cache with sync; echo 3 > /proc/sys/vm/drop_caches — needs root, a between-run reset, not a mid-training step.
G4: Thermal Throttling
  • SYMPTOM: throughput drops partway through a multi-hour run, or the box spontaneously reboots under sustained load.
  • CAUSE: sustained power draw caps around 100W versus the 240W rated figure; long runs push into that ceiling and throttle or, sometimes, reboot.
  • CHECK: references/gotcha-checks.md G4 — sample nvidia-smi --query-gpu=temperature.gpu,power.draw.
  • FIX: if power plateaus under 240W while temperature climbs, treat throttling as the cause; improve cooling or cap run length.
G5: Bandwidth Ceiling
  • SYMPTOM: memory-bound workloads, decode-heavy RL loops especially, plateau well below expected throughput.
  • CAUSE: 273 GB/s is a spec ceiling, not sustained; measured bandwidth runs 180–192 GB/s.
  • CHECK: references/gotcha-checks.md G5 — observed step time vs. the measured range, not spec.
  • FIX: budget throughput from 180–192 GB/s; revise a plan built on the 273 GB/s figure.
Show full SKILL.md (391 more words)Show less
G6: Global UMA Resource Contention
  • SYMPTOM: a process's KV cache/weights get evicted mid-run silently, no OOM in its own logs.
  • CAUSE: unified memory is one global pool; an uncapped or near-capacity process competes with anything else and can evict it. A small, bounded workload doesn't — a <4GB LoRA coexists fine alongside vLLM capped at gpu-memory-utilization<=0.5.
  • CHECK: references/gotcha-checks.md G6 — other GPU-resident processes and whether capped.
  • FIX: the one-heavy-job rule applies to uncapped or near-capacity workloads — cap or stop unrelated servers first. A small, capped workload need not stop.
G7: NVFP4 Slower Than FP8 on SM121
  • SYMPTOM: switching an inference workload from FP8 to NVFP4 on Spark makes it slower, not faster.
  • CAUSE: SM121 lacks cvt.e2m1x2 unless kernels target sm_121a; NVFP4 runs ~32% slower without it.
  • CHECK: references/gotcha-checks.md G7 — capability reports (12, 1); does the build target sm_121a?
  • FIX: stay on FP8 unless the build targets sm_121a.
G8: Stale Official Playbooks
  • SYMPTOM: following an official DGX Spark playbook still fails, with no local misconfiguration explaining it.
  • CAUSE: official playbooks have shipped broken before; the stack moves faster than the docs.
  • CHECK: references/gotcha-checks.md G8 — the playbook repo's recent issues.
  • FIX: check github.com/NVIDIA/dgx-spark-playbooks issues before trusting a recipe for an expensive run.
G9: Container-First, Not Bare Pip
  • SYMPTOM: a bare-pip environment that worked yesterday breaks after an unrelated pip install, or two "identical" environments behave differently.
  • CAUSE: bare pip lets Triton, xformers, and transformers drift independently; nothing pins them to GB10's SM121 target.
  • CHECK: references/gotcha-checks.md G9 — container or bare pip?
  • FIX: prefer an NGC container (see spark-environment-setup for tag guidance) or Unsloth's container. If bare pip is unavoidable, follow the NVIDIA install order, including --no-deps on Unsloth.
G10: Dual-Spark Is DDP/FSDP Only
  • SYMPTOM: a tensor-parallel launch across two Sparks hangs, runs far slower than single-Spark, or errors out.
  • CAUSE: ConnectX-7 is fast enough for gradient/parameter sync (DDP, FSDP) but too thin for TP's fine-grained traffic.
  • CHECK: references/gotcha-checks.md G10 — the configured parallelism strategy.
  • FIX: on a two-Spark setup, choose DDP or FSDP, never tensor parallelism — TP is single-node only here.

Fast Triage

The cheapest checks to run before anything else:

bash
python3 -c "import torch; print(torch.version.cuda)"  # expect 13.x (G1); NGC builds have no +cu130 tag — that's not a failure
python
import torch; print(torch.cuda.get_device_capability())  # expect (12, 1) (G7)
bash
{ [ -f /.dockerenv -o -f /run/.containerenv ] || grep -qE 'docker|containerd' /proc/1/cgroup; } 2>/dev/null && echo container || echo unknown  # G9

assets/preflight.sh runs G1, G3, G4, G7, G9 and produces one output line per gotcha in a fixed format: G-number first, then PASS/FAIL/WARN where automatable, SKIP when unavailable, or INFO: for a raw reading (G3, G4). Full commands: references/gotcha-checks.md. See also spark-environment-setup for the environment assumed working.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references, assets) in plugins/dgx-spark-ops/skills/spark-training-gotchas of wshobson/agents.

  • SKILL.md
  • assets/preflight.sh
  • references/gotcha-checks.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

DGX Spark Training Gotchas next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

DGX Spark Training Gotchas compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
DGX Spark Training Gotchas this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
GPU OptimizerMathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
Hyperpod Version Checkerawslabs/agent-plugins916—~910Automated safety check: PassApache-2.0
MUSA GPU Training Optimizeropen-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    329 GitHub stars~3.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check: notes
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    916 GitHub stars~910 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Mamba State-Space Models

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

    13k GitHub starsUsed in 2 repos~1.8k tokens
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 6 days ago
    Auto-check passed

Questions about DGX Spark Training Gotchas

What does DGX Spark Training Gotchas do?

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. The GB10 chip in DGX Spark, with Grace Blackwell, SM121, 128GB of unified memory and an aarch64 CPU, has ten recurring failure modes named G1 to G10 so tooling can check them by number. The skill is meant to be read before a long run, and applies when a run will not start, when it runs out of memory while `nvidia-smi` shows headroom, when throughput drops partway, before any multi-hour job, when linking two Sparks, or when choosing between FP8 and NVFP4.

When should I use DGX Spark Training Gotchas?

DGX Spark Training Gotchas fits situations like: starting a multi-hour training job on a DGX Spark; debugging an import error or segfault at launch on GB10; finding out why a run hit out-of-memory with free memory showing; choosing a parallelism strategy for two linked Sparks.

How do I install DGX Spark Training Gotchas in Claude Code?

Run `npx skills add wshobson/agents --skill spark-training-gotchas -a claude-code`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-training-gotchas in wshobson/agents) into .claude/skills/spark-training-gotchas in your project. Claude Code loads it when a task matches its description.

How do I install DGX Spark Training Gotchas in Codex?

Run `npx skills add wshobson/agents --skill spark-training-gotchas -a codex`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-training-gotchas in wshobson/agents) into .agents/skills/spark-training-gotchas in your project. Codex loads it when a task matches its description.

Can I use DGX Spark Training Gotchas in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill spark-training-gotchas -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-training-gotchas, .gemini/skills/spark-training-gotchas, .github/skills/spark-training-gotchas and .opencode/skills/spark-training-gotchas in your project.

What does DGX Spark Training Gotchas need to run?

Going by SKILL.md and its folder, DGX Spark Training Gotchas needs a shell for the scripts in its folder and the command-line tools its instructions call (pip and python3). Our summary lists: An NVIDIA DGX Spark (GB10) system; A shell to run assets/preflight.sh.

Does DGX Spark Training Gotchas access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is DGX Spark Training Gotchas safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does DGX Spark Training Gotchas use?

DGX Spark Training Gotchas is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does DGX Spark Training Gotchas use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to DGX Spark Training Gotchas?

Skills that share tags, products or a category with DGX Spark Training Gotchas: Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), GPU Optimizer (Mathews-Tom/armory, 329 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and Hyperpod Version Checker (awslabs/agent-plugins, 916 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains DGX Spark Training Gotchas?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.