Model Serving
ancoleman/ai-design-components
LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
$ npx skills add wshobson/agents --skill spark-memory-thermal-ops -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents spark-memory-thermal-ops --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops .claude/skills/spark-memory-thermal-ops && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "spark-memory-thermal-ops" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops into .claude/skills/spark-memory-thermal-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-memory-thermal-ops", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-opsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill spark-memory-thermal-ops -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents spark-memory-thermal-ops --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops .agents/skills/spark-memory-thermal-ops && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "spark-memory-thermal-ops" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops into .agents/skills/spark-memory-thermal-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-memory-thermal-ops", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill spark-memory-thermal-ops -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents spark-memory-thermal-ops --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops .cursor/skills/spark-memory-thermal-ops && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "spark-memory-thermal-ops" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops into .cursor/skills/spark-memory-thermal-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-memory-thermal-ops", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/dgx-spark-ops/skills/spark-memory-thermal-ops--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill spark-memory-thermal-ops -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents spark-memory-thermal-ops --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops .gemini/skills/spark-memory-thermal-ops && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "spark-memory-thermal-ops" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops into .gemini/skills/spark-memory-thermal-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-memory-thermal-ops", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents spark-memory-thermal-opsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill spark-memory-thermal-ops -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops .github/skills/spark-memory-thermal-ops && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "spark-memory-thermal-ops" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops into .github/skills/spark-memory-thermal-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-memory-thermal-ops", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill spark-memory-thermal-ops -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents spark-memory-thermal-ops --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops .opencode/skills/spark-memory-thermal-ops && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "spark-memory-thermal-ops" agent skill from https://github.com/wshobson/agents/tree/main/plugins/dgx-spark-ops/skills/spark-memory-thermal-ops into .opencode/skills/spark-memory-thermal-ops/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "spark-memory-thermal-ops", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
spark-memory-thermal-opsPlans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
DGX Spark's GB10 chip shares one 128GB unified memory pool between CPU and GPU, so habits from discrete GPUs mislead. nvidia-smi and cudaMemGetInfo can underreport pressure because page cache and mmap'd pages draw on the same pool, so the skill has the agent budget against free -g. It also notes that loading a model is a transient peak, since the mmapped weights and the CUDA copy count together for a while.
When a job runs out of memory, the agent works an OOM ladder in order: flush first, then change batch size or packing, then downgrade the method. For long runs the skill points to the sustained power ceiling and a thermal log, with assets/thermal-sample.sh, so a mid-job slowdown can be told apart from a configuration bug. A trainer and an inference server such as vLLM or Ollama should run one at a time. Launch-time failures belong to a sibling spark-training-gotchas skill, and references/uma-accounting.md holds the memory accounting.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell), which the agent can run.
Shell commands in SKILL.md call:
bashollamaFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
DGX Spark Memory and Thermal Ops loads about 2k tokens when it runs, and up to ~3.1k if it reads all its reference files. Until then it costs about 69 tokens; SKILL.md has 1,046 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,046 words, ~1,983 tokens.
.claude/skills/spark-memory-thermal-ops/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.DGX Spark's GB10 chip has one 128GB unified
memory (UMA) pool shared by CPU and GPU, and a
sustained power ceiling well below its rated
figure. Both break discrete-GPU assumptions:
headroom isn't what nvidia-smi reports, and a
run that starts fast will slow down mid-job
with nothing misconfigured. This skill covers
planning memory headroom, working an actual
OOM, and watching thermals across a long job.
For launch-time failure modes (ABI mismatches,
flash-attn, playbook breakage), see
spark-training-gotchas — this skill assumes
the job starts.
| Situation | Do this |
|---|---|
| Planning headroom before launch | Budget against free -g, not nvidia-smi — see UMA Memory Model |
| Job OOMs on unified memory | Work the OOM Ladder in order: flush, then batch/pack, then method downgrade |
| Throughput drops mid-run | Check the power/temp log before assuming a config bug — see Thermal Monitoring |
| Trainer + inference server both wanted | Run one at a time — see Concurrent Workloads |
Spark has no separate GPU VRAM — the GPU and CPU share one 128GB pool. Two consequences:
nvidia-smi and cudaMemGetInfo
underreport pressure — or report nothing at
all. Both report CUDA-allocator-visible
memory, not the pool's actual state — a box can
show headroom in nvidia-smi and still OOM,
because page-cache and mmap'd pages the
allocator doesn't see consume the same pool. On
some driver/setups, the memory query returns
[N/A], [N/A] outright instead of a number — a
script grepping for a numeric value there gets
nothing, not a misleading undercount (see
spark-training-gotchas gotcha G3).
Model load is a transient peak, not the steady state. Loading safetensors weights mmaps the file, then copies into CUDA tensors — for a window during load, both the mmap'd pages and the CUDA copy count against the pool at once. A model that fits while training can still OOM during load if headroom was sized for the post-load footprint instead of this doubled transient.
Plan and diagnose with free -g, not
nvidia-smi:
free -g | awk 'NR==2 {print "free:", $4, "GB"}'Rule of thumb: take that free figure, subtract a
few GB for OS/driver overhead, and budget against
the result — not the 128GB spec number.
The worksheet in references/uma-accounting.md
accepts parameter count, dtype, and method as
input, and returns a memory estimate to compare
against known anchors.
Before launch, work through these in order:
free -g; subtract OS/driver overhead
for the budget.references/uma-accounting.md.A sanity check of the worksheet formula against the ≈40GB anchor:
params = 70e9
weights_gb = params * 0.5 / 1e9 # NF4, step 1
adapter_gb = 0.5 # step 5, negligible
total_gb = weights_gb + adapter_gb # + activations
print(f"{total_gb:.0f}GB before activations")Weights alone land near the ≈40GB anchor — a plan estimating far above that for the same model class is a signal to recheck dtype and method.
When a job OOMs on unified memory, work this ladder in order. Each step is more disruptive than the last — don't skip ahead: reducing batch size is never step 1.
Flush the buffer cache. Page cache from a previous run or a large dataset read often accounts for GB of the "missing" headroom. This costs nothing but a rerun and doesn't touch the job's configuration:
sync; echo 3 > /proc/sys/vm/drop_cachesNeeds root; a between-run reset, not a
mid-training step. See
spark-training-gotchas (gotcha G3) for the
full diagnostic behind this step.
Reduce batch size or packing length. Only after a flush fails to free enough headroom, cut batch size or packing length — the first step that changes what the run does. Prefer packing length first; it drives activation footprint more directly at long context.
Downgrade the method: bf16 LoRA before QLoRA. If flushing and shrinking batch/pack still OOM, drop the method a tier — bf16 LoRA is next, not the reverse. QLoRA's bitsandbytes dequantization buffers are transient CUDA-side allocations that can OOM before an equivalent bf16 LoRA run would, even though QLoRA's steady-state footprint is smaller. A QLoRA OOM is not proof the model doesn't fit.
Fall back further (smaller model, multi-Spark) only after all three steps and the job still won't fit.
Multi-hour runs push into Spark's sustained power ceiling, well under the rated figure — expected platform behavior, not a symptom to explain away:
Sample temperature and power alongside the
training logs, not after a slowdown is
noticed — every 30-60 seconds correlates a
throughput drop with a thermal event. Keep
the CSV output format assets/thermal-sample.sh
writes, so timestamps line up against the log:
bash assets/thermal-sample.sh 30 thermal.logA sustained ~100W power draw is the platform cap, not a configuration bug. Don't re-tune batch size or precision to "fix" a plateau that's the box behaving normally under load. If temperature climbs while power stays flat under the rated 240W figure, that's the signature to recognize.
Log throttle events explicitly instead of
letting a run silently slow down unrecorded. A
run whose per-step time doubles two hours in
should show that in the log, correlated against
the thermal sample at that timestamp. Full
throttling diagnostics: spark-training-gotchas
(gotcha G4).
Because the 128GB pool is global, eviction happens without either process's logs showing an OOM:
The one-heavy-job rule applies to uncapped or
near-capacity workloads — an uncapped trainer
and inference server (vLLM, Ollama) compete for
the same pool. A small, capped workload doesn't:
a <4GB LoRA fine-tune coexists fine alongside
vLLM capped at gpu-memory-utilization<=0.5 —
check the other process's cap, not just its
presence, before stopping it.
Inference servers evict trainer pages silently under uncapped/near-capacity contention, and vice versa — neither logs an error, so a slow run or lost KV cache is a contention symptom to check for. Stop unrelated uncapped servers before a long or full-pool run.
Check for GPU-resident processes first:
ps aux | grep -E 'vllm|ollama|trl|axolotl' | grep -v grepThis procedure complements spark-training-gotchas
(gotchas G3, G4, G6) — that skill covers launch-time
failures; this one, the running job.
Memory math worksheets:
references/uma-accounting.md.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references, assets) in plugins/dgx-spark-ops/skills/spark-memory-thermal-ops of wshobson/agents.
Open the folder on GitHubat commit 46891e7
DGX Spark Memory and Thermal Ops next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| DGX Spark Memory and Thermal Ops this skillwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Model Servingancoleman/ai-design-components | 525 | — | ~3.4k | Automated safety check: Pass | MIT | |
| Jetson LLM BenchmarkNVIDIA/skills | 3.6k | 1 repos | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| GPU OptimizerMathews-Tom/armory | 329 | — | ~3.5k | Automated safety check: Notes | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.8k | Automated safety check: Pass | None |
ancoleman/ai-design-components
LLM and ML model deployment for inference. An agent skill from ancoleman/ai-design-components.
NVIDIA/skills
Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.
Mathews-Tom/armory
GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
K-Dense-AI/scientific-agent-skills
GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Implement distributed tracing with Jaeger and Tempo to track requests across microservices and identify performance bottlenecks.
wshobson/agents
Reference for designing and tuning production LLM prompts: few-shot examples, chain-of-thought, structured outputs, templates and system prompts.
Works with
Categories
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. DGX Spark's GB10 chip shares one 128GB unified memory pool between CPU and GPU, so habits from discrete GPUs mislead. nvidia-smi and cudaMemGetInfo can underreport pressure because page cache and mmap'd pages draw on the same pool, so the skill has the agent budget against free -g.
DGX Spark Memory and Thermal Ops fits situations like: sizing a training run against the 128GB pool before launch; recovering from an out-of-memory failure while loading or training on unified memory; logging temperature and power during a multi-hour job; deciding whether a mid-run slowdown is thermal throttling.
Run `npx skills add wshobson/agents --skill spark-memory-thermal-ops -a claude-code`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-memory-thermal-ops in wshobson/agents) into .claude/skills/spark-memory-thermal-ops in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill spark-memory-thermal-ops -a codex`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-memory-thermal-ops in wshobson/agents) into .agents/skills/spark-memory-thermal-ops in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill spark-memory-thermal-ops -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-memory-thermal-ops, .gemini/skills/spark-memory-thermal-ops, .github/skills/spark-memory-thermal-ops and .opencode/skills/spark-memory-thermal-ops in your project.
Going by SKILL.md and its folder, DGX Spark Memory and Thermal Ops needs a shell for the scripts in its folder and the command-line tools its instructions call (bash and ollama). Our summary lists: An NVIDIA DGX Spark with the GB10 chip; A shell with nvidia-smi and free -g for the checks.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
DGX Spark Memory and Thermal Ops is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with DGX Spark Memory and Thermal Ops: Model Serving (ancoleman/ai-design-components, 525 stars), Jetson LLM Benchmark (NVIDIA/skills, 3.6k stars), GPU Optimizer (Mathews-Tom/armory, 329 stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,314 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.