Hugging Face Local Model Evals
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results.
$ npx skills add NVIDIA/skills --skill tao-run-on-slurm -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills tao-run-on-slurm --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/tao-run-on-slurm .claude/skills/tao-run-on-slurm && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "tao-run-on-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-slurm into .claude/skills/tao-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-slurm", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-slurmType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill tao-run-on-slurm -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills tao-run-on-slurm --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/tao-run-on-slurm .agents/skills/tao-run-on-slurm && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "tao-run-on-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-slurm into .agents/skills/tao-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-slurm", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-run-on-slurm -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills tao-run-on-slurm --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/tao-run-on-slurm .cursor/skills/tao-run-on-slurm && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "tao-run-on-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-slurm into .cursor/skills/tao-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-slurm", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/tao-run-on-slurm--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill tao-run-on-slurm -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills tao-run-on-slurm --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/tao-run-on-slurm .gemini/skills/tao-run-on-slurm && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "tao-run-on-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-slurm into .gemini/skills/tao-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-slurm", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills tao-run-on-slurmInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill tao-run-on-slurm -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/tao-run-on-slurm .github/skills/tao-run-on-slurm && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "tao-run-on-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-slurm into .github/skills/tao-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-slurm", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill tao-run-on-slurm -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills tao-run-on-slurm --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/tao-run-on-slurm .opencode/skills/tao-run-on-slurm && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "tao-run-on-slurm" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/tao-run-on-slurm into .opencode/skills/tao-run-on-slurm/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "tao-run-on-slurm", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
tao-run-on-slurmRemote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results.
Tao Run On Slurm is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results. Use when running TAO training/eval/inference jobs on an on-prem or DGX SLURM cluster. Trigger phrases include "run on SLURM", "submit sbatch", "DGX SLURM cluster", "Pyxis/Enroot container", "Lustre dataset".
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including reference files (for example `BENCHMARK.md`, `config/skillspector-baseline.yaml` and `evals/evals.json`). Compatibility notes: Requires SSH access to a SLURM login node (passwordless via key auth) and SLURMUSER + SLURMHOSTNAME env vars. No nvidia-tao-sdk install is required; jobs are…
It sits in AI & LLM Engineering, covering GPU and accelerator computing. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadBashFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
sshbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use ssh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NGC_KEYHF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires SSH access to a SLURM login node (passwordless via key auth) and SLURM_USER + SLURM_HOSTNAME env vars. No nvidia-tao-sdk install is required; jobs are driven directly over ssh + sbatch/squeue/sacct/scancel.
From compatibility in the SKILL.md frontmatter.
Tao Run On Slurm loads about 4.8k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 85 tokens; SKILL.md has 2,145 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
set -a; source /path/to/.env; set +a # omit if already exportedallowed-tools: Read, BashAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 2,145 words, ~4,834 tokens.
.claude/skills/tao-run-on-slurm/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.Standalone install? If this session was not initialized by the TAO skill bank plugin, run the
tao-setupskill first (host preflight, credentials, cross-skill discovery).
Remote GPU compute platform for clusters managed by SLURM. Jobs are submitted
from the launch host to a login node over SSH, staged on a shared
filesystem, submitted with sbatch, and executed with srun container support.
Use SLURM when the user has access to a managed GPU cluster, shared Lustre storage, and scheduler-owned GPU allocation. Do not use SLURM for local files that exist only on the agent machine; data and outputs must be reachable from the cluster.
Confirm SLURM_USER and SLURM_HOSTNAME are exported and passwordless SSH to a
login host works (ssh -o BatchMode=yes).
The launch host needs ssh, not local sbatch, srun, Enroot, or a Lustre
mount. Preflight those scheduler, Pyxis, Enroot, and shared-storage dependencies
on the selected remote login/compute frame. Model-specific inspectors may be
streamed from the installed skill over SSH stdin; do not stage an ad-hoc source
patch or treat the launch host as the SLURM frame.
For private nvcr.io images, install ~/.config/enroot/.credentials on the
cluster once per (cluster, user): Pyxis/Enroot does not read NGC_KEY from the
job env, and without persistent credentials, auth-gated pulls fail with "Could
not process JSON input" at job startup. Install it via the printf | ssh
heredoc so the NGC_KEY value never lands in shell history, intermediate files,
or chat output; never cat/echo the value.
If a preflight check fails, the agent prompts the user to authorize the install/fix via Bash. Pip-installable Python requirements are the exception: install them automatically, then rerun preflight.
See references/slurm-ssh-credentials.md for the full preflight script, the
enroot-credentials heredoc, prerequisite key setup (keypair, ssh-copy-id,
known_hosts, container key mounts, 2FA handling), and the SSH failure
remediation prompt.
tao-run-on-slurm is a platform consumer: it runs a spec-bundle over
ssh + sbatch/squeue/sacct/scancel, mutating only the job-record. Storage is
tier A (Lustre) — the dataset is staged to a shared path before submit and
read through Pyxis; never fetch S3 inside the allocation (the scheduler-idle
timeout kills GPU-idle jobs and bills the wasted time). $BANK =
${TAO_SKILL_BANK_PATH}; $LOGIN = a resolved SLURM_HOSTNAME.
@@IMAGE@@ is a Lustre .sqsh — reuse an existing one if present
(ssh $LOGIN ls <sqsh>); only if missing, convert once with enroot import
(cached by name — see references/slurm-container-execution.md).ssh $LOGIN test -e …) and
reference those paths; tao-data-io stages only a small auxiliary input
that is not there yet — never re-stage existing data, and never the training
set inside the allocation.
Then author the spec at <job_dir>/specs/spec.yaml on Lustre with those paths.HF_TOKEN), write them to a mode-600 sidecar on Lustre and let the
template shred it on exit; NGC image pulls use the one-time
~/.config/enroot/.credentials (see references/slurm-ssh-credentials.md),
not the job env:set -a; source /path/to/.env; set +a # omit if already exported
printf 'export HF_TOKEN=%s\n' "$HF_TOKEN" | ssh $LOGIN "umask 077; cat > <job_dir>/job_$JOB_ID.env"results_dir on Lustre, before launch:JOB_ID=$("$BANK/scripts/tao_job_record.py" open --platform slurm --image "$IMAGE" \
--network-arch "$ARCH" --action "$ACTION" --storage-tier A --results-root "$SLURM_BASE_RESULTS_DIR")execution, preserve its order and semantics while mapping distributed
intent to native SLURM/Pyxis. Stage only its checksum-closed
supporting_files. The full generic lifecycle and staging contract is in
references/slurm-container-execution.md.templates/slurm/singlenode.sbatch.tmpl — substitute every
@@<NAME>@@ (JOB_NAME=$JOB_ID, NUM_GPUS, CPUS_PER_TASK, TIME, LOG_DIR,
IMAGE, CONTAINER_MOUNTS=<RUNTIME_SUPPLIED_MOUNTS>, COMMAND=<bundle command reading the shared-storage spec>, SBATCH_EXTRA= account/partition lines, ENV_FILE= the sidecar path or
empty, EXTRA_ENV= any cluster NCCL knobs) → <job_dir>/sbatch/job_$JOB_ID.sbatch.
Lint + syntax-check before submit: redact_secrets.py lint <sbatch> must
pass and bash -n <sbatch> must succeed.SLURM_ID=$(ssh $LOGIN "sbatch --parsable <job_dir>/sbatch/job_$JOB_ID.sbatch")
"$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state RUNNING --backend-ref "$SLURM_ID"A submit that skipped the gate or the open has no id — so it cannot launch.
# sacct ANNOTATES states ("CANCELLED by 12345") and truncates them to the
# default column width, so a cancelled job reads back as "CANCELLED+" and
# matches nothing in the table below — reporting UNKNOWN instead of CANCELED.
# Widen the column, take the first word, drop the truncation marker.
st=$(ssh $LOGIN "sacct -j $SLURM_ID -X -n -o State%30" | awk '{print $1}' | tr -d '+')
# (use squeue while the job is still PENDING; sacct lags briefly after submit)| SLURM state | vocab |
|---|---|
PENDING | PENDING |
RUNNING / COMPLETING | RUNNING |
COMPLETED | COMPLETE (confirm status.json in results_dir) |
FAILED / TIMEOUT / OUT_OF_MEMORY | ERROR (infra-vs-program classify → retry, M6) |
NODE_FAIL / BOOT_FAIL | ERROR, err_class=ERR_INFRA (--requeue re-queues these) |
CANCELLED / PREEMPTED / REVOKED | CANCELED |
| (not found) | UNKNOWN |
Native sub-state rides in the transition message. Poll at the chosen interval;
long queue waits are normal — do not stop on elapsed time.
ssh $LOGIN "tail -n ${N:-200} <log_dir>/$JOB_ID-$SLURM_ID/main.out" # SLURM auto-creates the %x-%j subdirssh $LOGIN "scancel $SLURM_ID"
"$BANK/scripts/tao_job_record.py" mark "$JOB_ID" --state CANCELED --source agentTreat an already-terminated SLURM job as a successful cancel.
Same four verbs, with three additions at submit:
templates/slurm/multinode.sbatch.tmpl instead of the single-node
one — it's a strict superset (adds --nodes / --wait-all-nodes + the
rendezvous block). WORLD_SIZE is the node count (TAO's misnomer); never
change it to a global-rank count.scripts/nccl_allreduce_probe.py under the container's torchrun) with a
~120s timeout. Before invoking torchrun, preserve the TAO rendezvous values
as TAO_NODE_COUNT=$WORLD_SIZE,
TAO_GPUS_PER_NODE=$NUM_GPU_PER_NODE, and
TAO_NODE_RANK=$SLURM_PROCID; torchrun overwrites its standard
WORLD_SIZE with the global process count. NCCL_PROBE_OK → proceed.
Timed out (the collective hung)
→ set the cluster's NCCL knob in EXTRA_ENV and re-probe — on CS-OCI-ORD that
is export NCCL_P2P_DISABLE=1 (the intra-node P2P hang), often with
NCCL_SOCKET_IFNAME=eth0 / NCCL_IB_DISABLE=1. Cache the working env per
cluster so later jobs skip the probe. Gate on gpus_per_node > 1 too —
the P2P hang triggers on a single node with 2+ GPUs.Read references/cosmos-slurm-guardrails.md
before rendering a Cosmos command. It defines image staging, planner
materialization, Framework and Cosmos-RL launch contracts, worker/runtime
requirements, and exit/status handling.
Use shared-filesystem URIs, not local or file:// paths; tao-core rejects
local/file paths for remote backends.
lustre:///absolute/path for user-provided datasets on Lustre.slurm:// paths may appear in microservices metadata and are converted to
Lustre paths before the container starts.Accept either dataset roots (model skills map them to required files) or direct
spec-key paths. After SSH succeeds and before generating scripts, test -e each
required dataset path from the login host; if it fails, stop and ask for
corrected paths or staged data rather than producing scripts that fail in the
first training job. See references/slurm-ssh-credentials.md for root vs.
direct-spec modes, backend details, and the results-dir default.
tao-core runs TAO containers through Pyxis/Enroot:
<job_dir>/specs, <job_dir>/env, and <job_dir>/meta.srun -n1 -p <conversion_partition> enroot import. This is a one-time cost
per image, not an optional optimization — see Acquire the image off the GPU
allocation below.<job_dir>/sbatch/job_<job_id>.sbatch.sbatch --export=ALL <script>.srun --container-image=<image> --container-mounts=<RUNTIME_SUPPLIED_MOUNTS>.Accepted image formats: /path/to/image.sqsh, registry#image:tag,
docker://registry#image:tag, and ordinary registry/image:tag (converted to
Pyxis form when needed). SQSH conversion is cached by image name; for :latest
images the cached SQSH is reused unless force_reconvert_latest is enabled.
The GPU is yours from the moment the allocation starts, not from when compute begins. Anything the job does before training — pulling a registry image, converting it, fetching a dataset — runs on GPUs that are idle, billed, and visible to the cluster's GPU-idle reaper. A first-time TAO pull plus enroot conversion is minutes of that, which is long enough to be killed and long enough to be expensive.
So the image must already be a local .sqsh when the GPU job starts. Passing a
docker:// or registry#image:tag URI straight to srun --container-image=
makes Pyxis pull and convert inside the allocation — the exact trap. Convert
once on a CPU partition, then point every later job at the resulting file:
# One-time per image, on CPU — costs no GPU time.
ssh $LOGIN "test -e <sqsh>" || \
ssh $LOGIN "srun --chdir=/tmp -n1 -c4 --mem=7200M \
-p <cpu_partition> -t <minutes> \
bash -c 'set -Eeuo pipefail
export TMPDIR=/tmp
export ENROOT_TEMP_PATH=/tmp/enroot-tao-\${SLURM_JOB_ID}
export SLURM_ENROOT_TEMP_PATH=\${ENROOT_TEMP_PATH}
mkdir -p \"\${ENROOT_TEMP_PATH}\"
cd /tmp
enroot import -o <sqsh> docker://<registry>#<image>:<tag>'"
# Every GPU job then references the file, never the registry.
srun --container-image=<sqsh> ...The same rule governs data: stage it to Lustre before submit (tier A) rather than fetching inside the allocation.
CS-OCI-ORD conversion uses cpu_long, 4 CPUs, 7200M memory, no exclusive node,
node-local Enroot temp paths, and at least 120 minutes. The execution reference
records the evidence and QOSGrpMemLimit recovery contract.
Partial conversions are self-detecting: the SQSH is validated by hsqs magic,
so a truncated file is rejected rather than silently used. Conversion runs once
and is then cached by image name.
A failed conversion must not fall back to the registry image. The tempting
recovery — pass docker://… to srun and let Pyxis handle it — puts the pull
back inside the GPU allocation, which is the cost the conversion existed to
avoid, and it does so precisely when something is already wrong. Treat a failed
or truncated conversion as fatal: fix it on the CPU partition and resubmit.
Diagnostic: if a job is unexpectedly slow to produce output, check what
--container-image= actually received. A registry URI there — rather than a
.sqsh path — means the pull happened on the GPUs.
squeue/sacct;
TAO terminal status comes from status.json in the shared results folder.PENDING, RUNNING, or otherwise). Do not stop after a
fixed elapsed time such as 30 minutes; long queue waits are normal on shared
GPU partitions.<job_dir>/slurm-logs/<slurm_job_name>-<slurm_job_id>/main.out and .err.backend_details.slurm_metadata.slurm_job_id and running
scancel <slurm_job_id> over SSH. Treat missing or already terminated jobs as
successful cancellation.Status mapping:
PENDING -> PendingRUNNING or COMPLETING -> RunningCOMPLETED -> check status.jsonFAILED, BOOT_FAIL, DEADLINE, OUT_OF_MEMORY, NODE_FAIL -> retry if
logs match retriable infrastructure patterns, otherwise ErrorCANCELLED, PREEMPTED, REVOKED -> CanceledTIMEOUT -> ErrorSUSPENDED, STOPPED -> Running (still scheduler-owned and may resume;
the native sub-state rides in the transition message — same convention as
docker paused)Ask for these in the SLURM intake; see references/slurm-ssh-credentials.md
for the full credential list, microservices schema keys, and defaults.
polar,polar3,polar4,grizzly, treated as 4-hour queues.SSH_AUTH_SOCK agent-socket fallback.#SBATCH --account.Do not ask for SLURM_ACCOUNT or SLURM_BASE_RESULTS_DIR in the initial
intake unless the user says their site requires an account, wants a custom
results root, or the workflow cannot proceed without overriding defaults.
Defaults from tao-core:
num_nodes: 1num_gpus: 4max_num_gpus_per_node: 8cpus_per_task: 16time_hours: 4timeout_hours: 3.8max_time_hours: 4container_mounts: explicit source-to-target mounts supplied at runtimeuse_requeue: trueuse_sqsh: trueLaunchers must use the packaged 4-hour wall and 3.8-hour child-timeout
defaults, never 12 hours. If the user supplies a longer
SLURM_TIME_HOURS, verify that the selected partition supports it before
submitting. For the packaged default partition list
polar,polar3,polar4,grizzly, reject requests above 4 hours and ask for a
different partition only if the user actually wants a longer wall time.
At or above max_num_gpus_per_node, allocate exclusive nodes and derive their
count from total GPUs.
For multi-node jobs (num_nodes > 1), the rendered
templates/slurm/multinode.sbatch.tmpl sets the sbatch directives and exports
the PyTorch-distributed rendezvous env vars: WORLD_SIZE, NUM_GPU_PER_NODE,
NODE_RANK, MASTER_ADDR, and MASTER_PORT (29500). TAO entrypoints read
WORLD_SIZE + NUM_GPU_PER_NODE and build torchrun internally. Cosmos-RL has
special multi-node role handling for controller, policy, and rollout workers.
See the ### Multi-node (nodes > 1) submit subsection above for the NCCL-probe
gate and per-cluster env caching.
Use Lustre, not S3, for SLURM job inputs. The GPU allocation starts the
moment the job is dispatched, so a long s3:// download at the top of the
script burns the allocation, can get the job killed for GPU-idle, and is billed
either way. Stage training data on the shared filesystem first and reference it
as lustre:///.... S3/HF/NGC pre-fetch is fine for small auxiliary inputs
(checkpoints, configs), not training datasets. K8s/Brev do not share this
scheduler-idle constraint.
On an infrastructure failure (NODE_FAIL, BOOT_FAIL, NCCL transport timeouts,
CUDA driver init failures, GPU/IB link-down, OOM-killer node reaping, Xid
errors), classify infra-vs-program from the logs and create a new retry record
with --retry-of before re-submitting the staged workload (M6). Plain training
failures surface immediately so a broken spec does not consume the retry
budget. #SBATCH --requeue is enabled by default via
SLURM_USE_REQUEUE=true, so SLURM itself re-queues the job on NODE_FAIL or
pre-emption before any agent-level resubmit; workload contracts such as Cosmos
may require --no-requeue.
Treat an empty sbatch --parsable response or SSH disconnect as ambiguous:
reconcile by exact job name, never submit blindly, and validate inherited node
exclusions. The referenced execution guide defines the full decision table.
See references/slurm-container-execution.md for the full multi-node
env-var/sbatch directive detail and table, cluster requirements, the
Lustre-not-S3 rule in full, and the failure-mode checklist.
references/slurm-ssh-credentials.md — preflight script, SSH/key setup,
enroot credentials, full credential list, backend details, storage rules,
SSH remediation prompt.references/slurm-container-execution.md — container execution steps,
monitoring, status mapping, cancellation, multi-node detail,
Lustre-not-S3, retries, failure modes.references/slurm-preflight-storage.md — extended preflight/storage notes.references/cosmos-slurm-guardrails.md — Cosmos Framework and Cosmos-RL
launch and status guardrails.references/detailed-guide.md — navigation map for the split references.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 11 other files (references) in skills/tao-run-on-slurm of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Tao Run On Slurm next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Tao Run On Slurm this skillNVIDIA/skills | 3.6k | — | ~4.8k | Automated safety check: Notes | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Liger Kernel Perflinkedin/Liger-Kernel | 6.7k | — | ~1.5k | Automated safety check: Pass | BSD-2-Clause | |
| Hugging Face LLM Trainerhuggingface/skills | 11k | 1 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | |
| MUSA GPU Training Optimizeropen-infra-skills/infra-skills | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| Cuda Kernel OptimizerKernelFlow-ops/cuda-optimized-skill | 214 | — | ~4.3k | Automated safety check: Pass | MIT |
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
linkedin/Liger-Kernel
Optimizes the performance of existing Liger Kernel Triton kernels.
huggingface/skills
Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.
open-infra-skills/infra-skills
Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.
KernelFlow-ops/cuda-optimized-skill
Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.
inclusionAI/AReno
Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Categories
Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results. Tao Run On Slurm is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results.
Tao Run On Slurm fits situations like: running TAO training/eval/inference jobs on an on-prem; DGX SLURM cluster; phrases include run on SLURM; pyxis/Enroot container.
Run `npx skills add NVIDIA/skills --skill tao-run-on-slurm -a claude-code`. Or copy the skill folder (skills/tao-run-on-slurm in NVIDIA/skills) into .claude/skills/tao-run-on-slurm in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill tao-run-on-slurm -a codex`. Or copy the skill folder (skills/tao-run-on-slurm in NVIDIA/skills) into .agents/skills/tao-run-on-slurm in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill tao-run-on-slurm -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tao-run-on-slurm, .gemini/skills/tao-run-on-slurm, .github/skills/tao-run-on-slurm and .opencode/skills/tao-run-on-slurm in your project.
Going by SKILL.md and its folder, Tao Run On Slurm needs the command-line tools its instructions call (ssh and bash) and credentials named NGC_KEY and HF_TOKEN. Our summary lists: Python 3; Docker; A credential in NGC_KEY. Its frontmatter pre-approves these tools: Read, Bash. Compatibility (from SKILL.md): Requires SSH access to a SLURM login node (passwordless via key auth) and SLURM_USER + SLURM_HOSTNAME env vars. No nvidia-tao-sdk install is required; jobs are driven directly over ssh + sbatch/squeue/sacct/scancel..
SKILL.md contains no URLs. Its commands use ssh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Tao Run On Slurm is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Tao Run On Slurm: Hugging Face Local Model Evals (huggingface/skills, 11k stars), Liger Kernel Perf (linkedin/Liger-Kernel, 6.7k stars), Hugging Face LLM Trainer (huggingface/skills, 11k stars) and MUSA GPU Training Optimizer (open-infra-skills/infra-skills, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.