Agent skill

Spark Environment Setup

by wshobson in wshobson/agents

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

MITAuto-check passedAI & LLM Engineering

Install Spark Environment Setup

skills CLI
$ npx skills add wshobson/agents --skill spark-environment-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents spark-environment-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/dgx-spark-ops/skills/spark-environment-setup .claude/skills/spark-environment-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spark-environment-setup
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
1,014 words
Files
3 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

  • Installing PyTorch/Unsloth/TRL/vLLM on DGX Spark
  • SKILL.md covers When to Use This Skill, Container-First Rule, The ABI Rule and Component Quick Table, plus 2 more sections
  • Calls pip, docker and python3
  • Hitting libcudart

What it does

Spark Environment Setup is an agent skill from wshobson/agents. Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). Use when installing PyTorch/Unsloth/TRL/vLLM on DGX Spark, hitting libcudart or wheel-ABI errors on aarch64, or choosing between NGC containers and bare pip installs.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/container-workflow.md` and `references/stack-matrix.md`).

It sits in AI & LLM Engineering, covering Deep learning, Fine-tuning and LLM inference and serving. It works with CUDA, PyTorch, NVIDIA AI Platform and vLLM. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Installing PyTorch/Unsloth/TRL/vLLM on DGX Spark
  • Hitting libcudart
  • Wheel-ABI errors on aarch64
  • Choosing between NGC containers and bare pip installs

Example prompts

  • “/spark-environment-setup”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • docker
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spark Environment Setup loads about 2k tokens when it runs, and up to ~4.8k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,014 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 1,014 words, ~2,017 tokens.

Download SKILL.mdSave it as .claude/skills/spark-environment-setup/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
spark-environment-setup
description
Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). Use when installing PyTorch/Unsloth/TRL/vLLM on DGX Spark, hitting libcudart or wheel-ABI errors on aarch64, or choosing between NGC containers and bare pip installs.

Spark Environment Setup

DGX Spark ships a GB10 Grace Blackwell chip: aarch64 CPU, SM121 GPU, 128GB unified memory, CUDA 13. This is a narrower and younger platform than a standard x86 CUDA 12 box, so package selection and ABI matching matter more than usual — the wheel ecosystem for aarch64 + CUDA 13 is still filling in.

When to Use This Skill

  • Setting up a fresh Spark box for training or inference.
  • Hitting an import error mentioning libcudart, a missing symbol, or a wheel that "installed fine but won't load."
  • A framework install (PyTorch, Unsloth, TRL, vLLM, xformers) fails, hangs, or silently falls back to CPU.
  • Deciding whether to use an NGC container or bare pip.
  • Restoring a working setup after an OS reinstall or a base-image update, needing to re-verify from scratch.

Each of these accepts the same general fix: match the container/wheel combination to CUDA 13 and SM121, don't fight the ABI.

Container-First Rule

Quick decision, before the detail below:

  • Standard training/inference work → NGC PyTorch container.
  • Unsloth-centric fine-tuning → Unsloth container (it ships the pinned Triton/xformers/transformers combination already validated for that path).
  • Neither fits (custom system package, local IDE interpreter) → bare pip, following the exact sequence further down.

Default to a container. Use nvcr.io/nvidia/pytorch:25.09-py3 as the base for general work — the newest tag confirmed working on this hardware; pull a newer blessed tag if locally available rather than hard-blocking on 25.11-py3. NGC's tag is dated, so running it directly is fine:

bash
docker run --runtime=nvidia --gpus all -it --rm \
  nvcr.io/nvidia/pytorch:25.09-py3

unsloth/unsloth:dgxspark-latest is a moving tag by contrast — resolve and pin its digest before running it for anything reproducible; the bare tag is a discovery step only, not the default invocation. Full pull-inspect-pin sequence and flag rationale/volume mounts for finetuning/ run dirs: references/container-workflow.md. Treat bare pip as the exception.

The reason for the container-first stance is pinning, not convenience. Triton, xformers, and transformers versions interact narrowly with GB10's SM121 target and CUDA 13; a container locks all of them together against a combination already validated on this hardware. Bare pip leaves that resolution to you, one broken import at a time.

When bare pip is warranted, follow the NVIDIA playbook's install sequence verbatim and in order:

bash
pip install "transformers==5.13.1" "peft==0.19.1" "hf_transfer==0.1.9" "datasets==4.3.0" "trl==1.8.0"
pip install --no-deps "unsloth==2026.7.2" "unsloth_zoo==2026.7.2" "bitsandbytes==0.49.2"
pip install -U "torchao==0.17.0"

The second command's --no-deps flag is not optional — letting pip re-resolve Unsloth's dependency tree on aarch64 is a common way to pull in an incompatible torch or triton build. The third line is not optional either: the NGC base image's bundled torchao is too old for current peft's LoRA-attach path (ImportError: ... torchao ... only versions above 0.16.0 are supported) — a hard blocker, not a warning. Every == pin above is load-bearing, taken from the dated known-good version matrix in references/stack-matrix.md (its Last verified date governs staleness) — an unpinned install resolves current PyPI versions well outside what this Unsloth release supports.

Pull a fresh tag when a new blessed release is announced. Rebuild locally from one of the two bases only when a project needs an extra system package layered in — not to "upgrade" a component the image already pins. Details on both paths: references/container-workflow.md.

One more preflight: official DGX Spark playbooks have shipped broken before. Check recent issues on github.com/NVIDIA/dgx-spark-playbooks (and the other resources in references/stack-matrix.md) before trusting a recipe verbatim for a long run.

Show full SKILL.md (486 more words)Show less

The ABI Rule

The single most common failure on Spark is a CUDA 12/13 ABI mismatch: a wheel built against libcudart.so.12 loaded on a system that only has libcudart.so.13. The install usually succeeds; the failure surfaces later as a missing-symbol error or a segfault that doesn't obviously point at CUDA.

Fix: pull wheels from download.pytorch.org/whl/cu130 (the cu130-tagged aarch64 builds), or use one of the containers above, which already carry a matched build. Before chasing a stack trace that mentions a CUDA symbol, check which CUDA tag the installed wheel was built against:

bash
python3 -c "import torch; print(torch.version.cuda)"

If that output doesn't start with 13, the ABI mismatch is the first thing to fix. NGC container builds (e.g. nvcr.io/nvidia/pytorch:25.09-py3) build torch internally against CUDA 13 with no +cu130 wheel tag — pip show torch won't say cu130 there, and that absence alone is not a failure.

Typical symptoms:

  • ImportError: undefined symbol referencing a CUDA runtime function.
  • A segfault on the first .cuda() call, no useful traceback.
  • A wheel that installs cleanly, then fails at import time — pip's resolver doesn't check CUDA ABI, only version constraints.
  • Two "identical" environments behaving differently — usually one has a cu130 wheel, the other a cu121/cu124 leftover.

The fix is the same regardless of symptom: match the wheel's CUDA tag to the system, or use a container that already does.

Component Quick Table

Condensed status for the components most likely to come up. Full table with wheel URLs, build flags, the sm_121 vs sm_121a distinction, and the dated known-good version matrix: references/stack-matrix.md.

ComponentStatus
PyTorch✅ official cu130 aarch64 wheels
bitsandbytes✅ works out of the box
Triton✅ needs the TRITON_PTXAS_PATH parameter set
flash-attn❌ skip pip build; NGC bundles a working one — see spark-training-gotchas G2
xformerssource build only (TORCH_CUDA_ARCH_LIST=12.1)
vLLMnightly wheels only
TransformerEngine / NVFP4 traincontainer-only

Everything else — Unsloth, Axolotl, TRL, PEFT — installs cleanly through the container-first path above. LLaMA-Factory and NeMo are fragile on Spark; check upstream issues first.

Verification Commands

Confirm the environment can actually see the GPU before running anything expensive:

python
import torch
print(torch.cuda.is_available(), torch.version.cuda)

This call returns two values; the exact output format is one line, <bool> <cuda-version>:

text
True 13.0

If it prints False instead, don't jump straight to a wheel reinstall — ABI mismatch is one cause among several:

HypothesisQuick check
Runtime/flagsnvidia-smi fails in-container too
Device visibilityecho $CUDA_VISIBLE_DEVICES
Permissionsls -l /dev/nvidia*
CUDA init statewedged process; retry fresh shell/container
ABI mismatch (usual culprit)torch.version.cuda not 13.x

Check nvidia-smi first — if it doesn't show the GPU, it's one of the first three, not ABI. Reinstall a wheel only once ABI is confirmed. Per-hypothesis detail: references/stack-matrix.md. Run right after the container starts, before installing project-specific packages.

One more check: if Triton kernel compilation fails once training starts, set TRITON_PTXAS_PATH=/usr/local/cuda/bin/ptxas and retry — see references/stack-matrix.md for the full workaround list.

Next Steps

A verified environment is only the starting point. See also: spark-training-gotchas for failure preflights before a training run, and spark-memory-thermal-ops for unified-memory OOMs and thermal throttling during long ones.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/dgx-spark-ops/skills/spark-environment-setup of wshobson/agents.

  • SKILL.md
  • references/container-workflow.md
  • references/stack-matrix.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Spark Environment Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spark Environment Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spark Environment Setup this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Quark Env Preflightamd/Quark181—~1.4kAutomated safety check: PassMIT
Magpie Kernel Evaluatoramd/skills398—~2.3kAutomated safety check: PassMIT
Jetson PackageNVIDIA/skills3.5k1 repos~1.8kAutomated safety check: PassApache-2.0
Torch TensorrtVectorSpaceLab/AREX-Skill328—~1.5kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    398 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Jetson Package

    NVIDIA/skills

    Official

    Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

    3.5k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Torch Tensorrt

    VectorSpaceLab/AREX-Skill

    A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging…

    328 GitHub stars~1.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 14 repos~473 tokens
    Auto-check passed
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 13 repos~502 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Questions about Spark Environment Setup

What does Spark Environment Setup do?

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). Spark Environment Setup is an agent skill from wshobson/agents. Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

When should I use Spark Environment Setup?

Spark Environment Setup fits situations like: installing PyTorch/Unsloth/TRL/vLLM on DGX Spark; hitting libcudart; wheel-ABI errors on aarch64; choosing between NGC containers and bare pip installs.

How do I install Spark Environment Setup in Claude Code?

Run `npx skills add wshobson/agents --skill spark-environment-setup -a claude-code`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-environment-setup in wshobson/agents) into .claude/skills/spark-environment-setup in your project. Claude Code loads it when a task matches its description.

How do I install Spark Environment Setup in Codex?

Run `npx skills add wshobson/agents --skill spark-environment-setup -a codex`. Or copy the skill folder (plugins/dgx-spark-ops/skills/spark-environment-setup in wshobson/agents) into .agents/skills/spark-environment-setup in your project. Codex loads it when a task matches its description.

Can I use Spark Environment Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill spark-environment-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spark-environment-setup, .gemini/skills/spark-environment-setup, .github/skills/spark-environment-setup and .opencode/skills/spark-environment-setup in your project.

What does Spark Environment Setup need to run?

Going by SKILL.md and its folder, Spark Environment Setup needs the command-line tools its instructions call (pip, docker and python3). Our summary lists: Python 3; Docker.

Does Spark Environment Setup access the network?

SKILL.md contains no URLs. Its commands use pip and docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Spark Environment Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Spark Environment Setup use?

Spark Environment Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spark Environment Setup use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.8k tokens, read only when the agent opens those files.

What are the alternatives to Spark Environment Setup?

Skills that share tags, products or a category with Spark Environment Setup: Graphsignal (graphsignal/graphsignal, 257 stars), Quark Env Preflight (amd/Quark, 181 stars), Magpie Kernel Evaluator (amd/skills, 398 stars) and Jetson Package (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spark Environment Setup?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.