Agent skill

DGX Spark Memory and Thermal Ops

by wshobson in wshobson/agents

Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

MITAuto-check passedAI & LLM Engineering

Install DGX Spark Memory and Thermal Ops

skills CLI
$ npx skills add wshobson/agents --skill spark-memory-thermal-ops -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents spark-memory-thermal-ops --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops .claude/skills/spark-memory-thermal-ops && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spark-memory-thermal-ops
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
1,046 words
Files
3 (incl. references, assets)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

  • Works in 4 steps: Read free -g; subtract OS/driver overhead → Estimate weights + optimizer + gradients + → Compare against the closest anchor (70B → …
  • Sizing a training run against the 128GB pool before launch
  • SKILL.md covers Common Issues Quick Reference, When to Use This Skill, UMA Memory Model and The OOM Ladder, plus 2 more sections
  • Runs Shell scripts from its folder; calls bash and ollama

What it does

DGX Spark's GB10 chip shares one 128GB unified memory pool between CPU and GPU, so habits from discrete GPUs mislead. nvidia-smi and cudaMemGetInfo can underreport pressure because page cache and mmap'd pages draw on the same pool, so the skill has the agent budget against free -g. It also notes that loading a model is a transient peak, since the mmapped weights and the CUDA copy count together for a while.

When a job runs out of memory, the agent works an OOM ladder in order: flush first, then change batch size or packing, then downgrade the method. For long runs the skill points to the sustained power ceiling and a thermal log, with assets/thermal-sample.sh, so a mid-job slowdown can be told apart from a configuration bug. A trainer and an inference server such as vLLM or Ollama should run one at a time. Launch-time failures belong to a sibling spark-training-gotchas skill, and references/uma-accounting.md holds the memory accounting.

When your agent uses it

  • Sizing a training run against the 128GB pool before launch
  • Recovering from an out-of-memory failure while loading or training on unified memory
  • Logging temperature and power during a multi-hour job
  • Deciding whether a mid-run slowdown is thermal throttling

Example prompts

  • “Work out whether my fine-tune will fit in the Spark's unified memory before I launch it.”
  • “My training run died with an OOM even though nvidia-smi showed free memory, so tell me what to try first.”
  • “Set up temperature and power logging for a multi-hour run and flag when it throttles.”

Requirements

  • An NVIDIA DGX Spark with the GB10 chip
  • A shell with nvidia-smi and free -g for the checks

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read free -g; subtract OS/driver overhead
  2. Estimate weights + optimizer + gradients +
  3. Compare against the closest anchor (70B
  4. If the estimate is close to the budget, start

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • ollama

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

DGX Spark Memory and Thermal Ops loads about 2k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 1,046 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,046 words, ~1,983 tokens.

Download SKILL.mdSave it as .claude/skills/spark-memory-thermal-ops/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
spark-memory-thermal-ops
description
Manage unified memory and thermals during long-running ML jobs on NVIDIA DGX Spark. Use when planning memory headroom for a training run on GB10, when a job OOMs on unified memory, or when monitoring temperature and power during multi-hour training.

Spark Memory & Thermal Ops

DGX Spark's GB10 chip has one 128GB unified memory (UMA) pool shared by CPU and GPU, and a sustained power ceiling well below its rated figure. Both break discrete-GPU assumptions: headroom isn't what nvidia-smi reports, and a run that starts fast will slow down mid-job with nothing misconfigured. This skill covers planning memory headroom, working an actual OOM, and watching thermals across a long job. For launch-time failure modes (ABI mismatches, flash-attn, playbook breakage), see spark-training-gotchas — this skill assumes the job starts.

Common Issues Quick Reference

SituationDo this
Planning headroom before launchBudget against free -g, not nvidia-smi — see UMA Memory Model
Job OOMs on unified memoryWork the OOM Ladder in order: flush, then batch/pack, then method downgrade
Throughput drops mid-runCheck the power/temp log before assuming a config bug — see Thermal Monitoring
Trainer + inference server both wantedRun one at a time — see Concurrent Workloads

When to Use This Skill

  • Sizing a training run against the 128GB pool before launch — will this model, method, and batch/pack combination fit.
  • A run OOMs mid-load or mid-step and the remediation order matters — what to try first, second, third.
  • Watching temperature and power during a multi-hour job, deciding whether a slowdown is thermal throttling or something else.
  • Planning to run a trainer alongside an inference server (vLLM, Ollama) on the same box.

UMA Memory Model

Spark has no separate GPU VRAM — the GPU and CPU share one 128GB pool. Two consequences:

  • nvidia-smi and cudaMemGetInfo underreport pressure — or report nothing at all. Both report CUDA-allocator-visible memory, not the pool's actual state — a box can show headroom in nvidia-smi and still OOM, because page-cache and mmap'd pages the allocator doesn't see consume the same pool. On some driver/setups, the memory query returns [N/A], [N/A] outright instead of a number — a script grepping for a numeric value there gets nothing, not a misleading undercount (see spark-training-gotchas gotcha G3).

  • Model load is a transient peak, not the steady state. Loading safetensors weights mmaps the file, then copies into CUDA tensors — for a window during load, both the mmap'd pages and the CUDA copy count against the pool at once. A model that fits while training can still OOM during load if headroom was sized for the post-load footprint instead of this doubled transient.

Plan and diagnose with free -g, not nvidia-smi:

bash
free -g | awk 'NR==2 {print "free:", $4, "GB"}'

Rule of thumb: take that free figure, subtract a few GB for OS/driver overhead, and budget against the result — not the 128GB spec number. The worksheet in references/uma-accounting.md accepts parameter count, dtype, and method as input, and returns a memory estimate to compare against known anchors.

Planning Sequence

Before launch, work through these in order:

  1. Read free -g; subtract OS/driver overhead for the budget.
  2. Estimate weights + optimizer + gradients + activations from references/uma-accounting.md.
  3. Compare against the closest anchor (70B QLoRA, 27B LoRA, 9B full FT), not the estimate alone.
  4. If the estimate is close to the budget, start with shorter packing or a smaller batch — cheaper than hitting the OOM Ladder mid-run.
Example: Sizing a 70B QLoRA Run

A sanity check of the worksheet formula against the ≈40GB anchor:

python
params = 70e9
weights_gb = params * 0.5 / 1e9      # NF4, step 1
adapter_gb = 0.5                     # step 5, negligible
total_gb = weights_gb + adapter_gb   # + activations
print(f"{total_gb:.0f}GB before activations")

Weights alone land near the ≈40GB anchor — a plan estimating far above that for the same model class is a signal to recheck dtype and method.

Show full SKILL.md (509 more words)Show less

The OOM Ladder

When a job OOMs on unified memory, work this ladder in order. Each step is more disruptive than the last — don't skip ahead: reducing batch size is never step 1.

  1. Flush the buffer cache. Page cache from a previous run or a large dataset read often accounts for GB of the "missing" headroom. This costs nothing but a rerun and doesn't touch the job's configuration:

    bash
    sync; echo 3 > /proc/sys/vm/drop_caches

    Needs root; a between-run reset, not a mid-training step. See spark-training-gotchas (gotcha G3) for the full diagnostic behind this step.

  2. Reduce batch size or packing length. Only after a flush fails to free enough headroom, cut batch size or packing length — the first step that changes what the run does. Prefer packing length first; it drives activation footprint more directly at long context.

  3. Downgrade the method: bf16 LoRA before QLoRA. If flushing and shrinking batch/pack still OOM, drop the method a tier — bf16 LoRA is next, not the reverse. QLoRA's bitsandbytes dequantization buffers are transient CUDA-side allocations that can OOM before an equivalent bf16 LoRA run would, even though QLoRA's steady-state footprint is smaller. A QLoRA OOM is not proof the model doesn't fit.

Fall back further (smaller model, multi-Spark) only after all three steps and the job still won't fit.

Thermal Monitoring

Multi-hour runs push into Spark's sustained power ceiling, well under the rated figure — expected platform behavior, not a symptom to explain away:

  • Sample temperature and power alongside the training logs, not after a slowdown is noticed — every 30-60 seconds correlates a throughput drop with a thermal event. Keep the CSV output format assets/thermal-sample.sh writes, so timestamps line up against the log:

    bash
    bash assets/thermal-sample.sh 30 thermal.log
  • A sustained ~100W power draw is the platform cap, not a configuration bug. Don't re-tune batch size or precision to "fix" a plateau that's the box behaving normally under load. If temperature climbs while power stays flat under the rated 240W figure, that's the signature to recognize.

  • Log throttle events explicitly instead of letting a run silently slow down unrecorded. A run whose per-step time doubles two hours in should show that in the log, correlated against the thermal sample at that timestamp. Full throttling diagnostics: spark-training-gotchas (gotcha G4).

Concurrent Workloads

Because the 128GB pool is global, eviction happens without either process's logs showing an OOM:

  • The one-heavy-job rule applies to uncapped or near-capacity workloads — an uncapped trainer and inference server (vLLM, Ollama) compete for the same pool. A small, capped workload doesn't: a <4GB LoRA fine-tune coexists fine alongside vLLM capped at gpu-memory-utilization<=0.5 — check the other process's cap, not just its presence, before stopping it.

  • Inference servers evict trainer pages silently under uncapped/near-capacity contention, and vice versa — neither logs an error, so a slow run or lost KV cache is a contention symptom to check for. Stop unrelated uncapped servers before a long or full-pool run.

Check for GPU-resident processes first:

bash
ps aux | grep -E 'vllm|ollama|trl|axolotl' | grep -v grep

This procedure complements spark-training-gotchas (gotchas G3, G4, G6) — that skill covers launch-time failures; this one, the running job.

Memory math worksheets: references/uma-accounting.md.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references, assets) in plugins/dgx-spark-ops/skills/spark-memory-thermal-ops of wshobson/agents.

  • SKILL.md
  • assets/thermal-sample.sh
  • references/uma-accounting.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

DGX Spark Memory and Thermal Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

DGX Spark Memory and Thermal Ops compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
DGX Spark Memory and Thermal Ops this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Model Servingancoleman/ai-design-components525—~3.4kAutomated safety check: PassMIT
Jetson LLM BenchmarkNVIDIA/skills3.6k1 repos~3.1kAutomated safety check: PassApache-2.0
GPU OptimizerMathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNone

Similar skills

  • Model Serving

    ancoleman/ai-design-components

    LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.

    525 GitHub stars~3.4k tokensUpdated 10 mo ago
    AI & LLM EngineeringAuto-check passed
  • Official

    Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

    3.6k GitHub starsUsed in 1 repo~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    329 GitHub stars~3.5k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    938 GitHub stars~2.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize For GPU

    K-Dense-AI/scientific-agent-skills

    GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Data & AnalyticsAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 13 repos~473 tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Distributed Tracing

    wshobson/agents

    Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.

    40k GitHub starsUsed in 12 repos~527 tokens
    Auto-check passed
  • Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.

    40k GitHub stars~1.3k tokensUpdated 5 days ago
    Auto-check passed

Questions about DGX Spark Memory and Thermal Ops

What does DGX Spark Memory and Thermal Ops do?

Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. DGX Spark's GB10 chip shares one 128GB unified memory pool between CPU and GPU, so habits from discrete GPUs mislead. nvidia-smi and cudaMemGetInfo can underreport pressure because page cache and mmap'd pages draw on the same pool, so the skill has the agent budget against free -g.

When should I use DGX Spark Memory and Thermal Ops?

DGX Spark Memory and Thermal Ops fits situations like: sizing a training run against the 128GB pool before launch; recovering from an out-of-memory failure while loading or training on unified memory; logging temperature and power during a multi-hour job; deciding whether a mid-run slowdown is thermal throttling.

How do I install DGX Spark Memory and Thermal Ops in Claude Code?

Run `npx skills add wshobson/agents --skill spark-memory-thermal-ops -a claude-code`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-memory-thermal-ops in wshobson/agents) into .claude/skills/spark-memory-thermal-ops in your project. Claude Code loads it when a task matches its description.

How do I install DGX Spark Memory and Thermal Ops in Codex?

Run `npx skills add wshobson/agents --skill spark-memory-thermal-ops -a codex`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-memory-thermal-ops in wshobson/agents) into .agents/skills/spark-memory-thermal-ops in your project. Codex loads it when a task matches its description.

Can I use DGX Spark Memory and Thermal Ops in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill spark-memory-thermal-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-memory-thermal-ops, .gemini/skills/spark-memory-thermal-ops, .github/skills/spark-memory-thermal-ops and .opencode/skills/spark-memory-thermal-ops in your project.

What does DGX Spark Memory and Thermal Ops need to run?

Going by SKILL.md and its folder, DGX Spark Memory and Thermal Ops needs a shell for the scripts in its folder and the command-line tools its instructions call (bash and ollama). Our summary lists: An NVIDIA DGX Spark with the GB10 chip; A shell with nvidia-smi and free -g for the checks.

Does DGX Spark Memory and Thermal Ops access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is DGX Spark Memory and Thermal Ops safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does DGX Spark Memory and Thermal Ops use?

DGX Spark Memory and Thermal Ops is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does DGX Spark Memory and Thermal Ops use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.

What are the alternatives to DGX Spark Memory and Thermal Ops?

Skills that share tags, products or a category with DGX Spark Memory and Thermal Ops: Model Serving (ancoleman/ai-design-components, 525 stars), Jetson LLM Benchmark (NVIDIA/skills, 3.6k stars), GPU Optimizer (Mathews-Tom/armory, 329 stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains DGX Spark Memory and Thermal Ops?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.