Official agent skill

Nemo Automodel Launcher Config

by NVIDIA in NVIDIA/skills

Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.

OfficialApache-2.0Auto-check passed

Install Nemo Automodel Launcher Config

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-automodel-launcher-config -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-automodel-launcher-config --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-automodel-launcher-config .claude/skills/nemo-automodel-launcher-config && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-automodel-launcher-config
GitHub stars
3.6k
Token cost
~2.2k tokens
SKILL.md length
889 words
Files
5
Skills in repo
390
Repo updated
First seen
Licence
Apache-2.0

At a glance

Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.

  • Works in 3 steps: Interactive (default): runs torchrun on… → Slurm: submits a batch job to an HPC… → SkyPilot: cloud-agnostic job submission…
  • SKILL.md covers Instructions, Routing Boundary, Launch Methods and Interactive Launch, plus 6 more sections
  • Needs HF_TOKEN and WANDB_API_KEY

What it does

Nemo Automodel Launcher Config is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `BENCHMARK.md`, `evals/evals.json` and `skill-card.md`).

The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

Example prompts

  • “/nemo-automodel-launcher-config”

Requirements

  • Python 3
  • A credential in WANDB_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Interactive (default): runs torchrun on the current node. Suitable for single-node development and debugging.
  2. Slurm: submits a batch job to an HPC cluster scheduler. Handles multi-node setup, container management, and environment configuration.
  3. SkyPilot: cloud-agnostic job submission to AWS, GCP, Azure, Lambda, or Kubernetes. Supports spot instances.

What it can do on your machine

Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • WANDB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Automodel Launcher Config loads about 2.2k tokens when it runs. Until then it costs about 34 tokens; SKILL.md has 889 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~34
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 889 words, ~2,200 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-automodel-launcher-config/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
nemo-automodel-launcher-config
description
Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.
when_to_use
Configuring Slurm or SkyPilot job submission, setting up multi-node launch scripts, debugging job submission failures, or switching between interactive and…
license
Apache-2.0
metadata.author
NVIDIA
metadata.tags
nemo-automodel, launcher-config

Launcher Configuration

NeMo AutoModel supports three launch methods: interactive (torchrun), Slurm (HPC clusters), and SkyPilot (cloud-agnostic).

Instructions

For launcher questions, answer directly from this skill without inspecting the repository unless the user asks you to edit files. Keep the answer focused on the relevant launch YAML, required fields, and the expected runtime behavior.

Use these compact answer patterns for common questions:

  • Slurm multi-node: show a slurm: YAML block with job_name, nodes, ntasks_per_node, time, account or partition, container_image, hf_home, optional extra_mounts, env_vars, and master_port; explain that the launcher derives WORLD_SIZE = nodes * ntasks_per_node and sets MASTER_ADDR and MASTER_PORT.
  • SkyPilot spot: show a skypilot: YAML block with cloud, accelerators, num_nodes, use_spot: true, disk_size, region, setup, and env_vars; warn that spot instances can be preempted, set a short step_scheduler.checkpoint_interval, and resume with restore_from.path.
  • Nsight Systems on Slurm: show slurm.nsys_enabled: true alongside normal Slurm fields, say the launcher wraps the training command with nsys profile, and state that it produces a .nsys-rep report file. Treat profiling as diagnostic-only: use short profiling runs and disable it for normal production training because it adds overhead and large artifacts.

For Slurm answers, start with this minimal template and then adjust only the fields the user asked about:

yaml
slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  master_port: 13742
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"

For Slurm-only questions, do not discuss SkyPilot or profiling unless the user asks. For profiling questions, say the .nsys-rep report is written in the Slurm job working or output directory, using the launcher's Nsys output setting when one is configured.

Routing Boundary

Use this skill only for launch mechanics: interactive execution, Slurm, SkyPilot, containers, mounts, environment variables, rendezvous settings, and profiling.

Do not use this skill for implementing or registering new model architectures, Hugging Face state-dict adapters, model files, or capability flags. Those are model onboarding tasks, not launcher configuration tasks.

Launch Methods

  1. Interactive (default): runs torchrun on the current node. Suitable for single-node development and debugging.
  2. Slurm: submits a batch job to an HPC cluster scheduler. Handles multi-node setup, container management, and environment configuration.
  3. SkyPilot: cloud-agnostic job submission to AWS, GCP, Azure, Lambda, or Kubernetes. Supports spot instances.

Interactive Launch

bash
# Single GPU
automodel finetune llm -c config.yaml

# Multi-GPU (all GPUs on current node)
torchrun --nproc_per_node=8 -m nemo_automodel._cli.app finetune llm -c config.yaml

No additional YAML section is needed for interactive mode. The CLI routes to torchrun automatically when no slurm: or skypilot: section is present in the config.

Slurm Configuration

The SlurmConfig dataclass generates an SBATCH script from a template.

YAML Example
yaml
slurm:
  job_name: llm_finetune
  nodes: 2
  ntasks_per_node: 8
  time: "04:00:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  hf_home: ~/.cache/huggingface
  extra_mounts:
    - source: /data
      dest: /data
  env_vars:
    WANDB_API_KEY: "${WANDB_API_KEY}"
    HF_TOKEN: "${HF_TOKEN}"
Key Fields
  • job_name: Slurm job identifier
  • nodes: number of nodes to request
  • ntasks_per_node: number of tasks (GPUs) per node
  • time: wall-time limit in HH:MM:SS format
  • account, partition: Slurm scheduling parameters
  • container_image: Enroot/Pyxis container image path
  • nemo_mount: mount point for NeMo AutoModel source inside the container
  • hf_home: HuggingFace cache directory path
  • extra_mounts: list of VolumeMapping(source, dest) for additional container bind mounts
  • master_port: port for distributed communication (default 13742)
  • env_vars: environment variables passed into the job
  • nsys_enabled: when true, wraps the training command with nsys profile for Nsight Systems profiling

SkyPilot Configuration

The SkyPilotConfig dataclass defines cloud job parameters.

YAML Example
yaml
skypilot:
  cloud: aws
  accelerators: "H100:8"
  num_nodes: 2
  use_spot: true
  disk_size: 200
  region: us-east-1
  setup: "pip install nemo-automodel"
  env_vars:
    HF_TOKEN: "${HF_TOKEN}"
Key Fields
  • cloud: target cloud provider (aws, gcp, azure, lambda, kubernetes)
  • accelerators: GPU type and count (e.g., "H100:8", "A100-80GB:4")
  • num_nodes: number of cloud instances
  • use_spot: use preemptible/spot instances for cost savings
  • disk_size: disk size in GB per node
  • region: cloud region for instance placement
  • setup: shell commands to run before the training job (e.g., install dependencies)
  • env_vars: environment variables for the job
Show full SKILL.md (345 more words)Show less
SkyPilot spot checklist

When using spot or preemptible instances:

  • Set use_spot: true in the skypilot: section.
  • Include accelerators, num_nodes, disk_size, region, setup, and required env_vars.
  • Use short checkpoint intervals in the recipe, for example step_scheduler.checkpoint_interval, because spot instances can be preempted.
  • Resume from the most recent checkpoint after preemption with the recipe's restore_from setting.

Minimal spot-resume recipe keys:

yaml
step_scheduler:
  checkpoint_interval: 100

restore_from:
  path: /checkpoints/latest

Multi-Node Environment

For multi-node training (both Slurm and SkyPilot), the launcher automatically configures:

  • MASTER_ADDR: hostname of the first node
  • MASTER_PORT: port for rendezvous (default 13742)
  • WORLD_SIZE: total number of processes (nodes * ntasks_per_node)
  • NCCL environment variables for optimized collective communication

Nsys Profiling

Enable Nsight Systems profiling in Slurm jobs:

yaml
slurm:
  job_name: llm_profile
  nodes: 1
  ntasks_per_node: 8
  time: "00:30:00"
  account: my_account
  partition: batch
  container_image: nvcr.io/nvidia/nemo:dev
  nsys_enabled: true

This is a Slurm launcher setting. Normal Slurm fields such as job_name, nodes, ntasks_per_node, time, account or partition, and container_image still apply.

When nsys_enabled: true, the launcher wraps the training command with nsys profile and writes a .nsys-rep report file for performance analysis in the Slurm job working or output directory. Profiling is diagnostic-only: run it for a short investigation, expect overhead and large artifacts, and turn it off for normal production training.

Code Anchors

  • components/launcher/slurm/config.py - SlurmConfig dataclass, VolumeMapping
  • components/launcher/slurm/template.py - SBATCH script template generation
  • components/launcher/slurm/utils.py - Slurm submission utilities
  • components/launcher/skypilot/config.py - SkyPilotConfig dataclass
  • _cli/app.py - CLI entry point and launcher routing logic

Pitfalls

  • Port collisions: if the default master_port (13742) is in use by another job on the same node, change it to avoid connection failures.
  • Container mounts: the source path in extra_mounts must exist on all nodes in the allocation. Missing paths cause container startup failures.
  • Slurm fault tolerance: the fault tolerance plugin is Slurm-specific and does not work with SkyPilot or interactive mode.
  • SkyPilot spot preemption: spot instances (use_spot: true) may be preempted by the cloud provider. Enable checkpointing with short intervals to minimize lost work.
  • Environment variable syntax: use ${VAR} syntax in YAML for shell variable expansion. Bare variable names will not be expanded.
  • Time limit vs async checkpoint: if the Slurm time limit is too short, an in-progress async checkpoint write may be killed before completion, resulting in a corrupted checkpoint. Leave at least 5-10 minutes of margin.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/nemo-automodel-launcher-config of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 14a98ae

Compare with similar skills

Nemo Automodel Launcher Config next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Automodel Launcher Config compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Automodel Launcher Config this skillNVIDIA/skills3.6k—~2.2kAutomated safety check: PassApache-2.0
Configure Channelopenclaw/openclaw392k—~946Automated safety check: PassMIT
Shipping and Launch Checklistaddyosmani/agent-skills105k1 repos~2.8kAutomated safety check: PassMIT
Slurm Job Script GeneratorFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~1.3kAutomated safety check: NotesNone
Technical Job Searchgithub/awesome-copilot40k—~1.2kAutomated safety check: PassMIT
Launch With Slurmmlc-ai/pith-train355—~1.4kAutomated safety check: PassApache-2.0

Similar skills

  • Configure Channel

    openclaw/openclaw

    Configure and prove a chat channel with non-interactive one-liners; secrets only as SecretRefs.

    392k GitHub stars~946 tokensUpdated today
    Auto-check passed
  • Shipping and Launch Checklist

    addyosmani/agent-skills

    Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.

    105k GitHub starsUsed in 1 repo~2.8k tokens
    DevOps & CloudAuto-check passed
  • Slurm Job Script Generator

    FreedomIntelligence/OpenClaw-Medical-Skills

    Generate SLURM sbatch job scripts and sanity-check HPC resource requests (nodes, tasks, CPUs, memory, GPUs) for simulation runs.

    3.1k GitHub starsUsed in 1 repo~1.3k tokens
    DevelopmentAuto-check: notes
  • Technical Job Search

    github/awesome-copilot

    Official

    A skill your agent uses when a software engineer asks for help with job search tasks: parsing or analyzing a job description, tailoring a CV/resume, writing a cover letter, evaluating a job offer…

    40k GitHub stars~1.2k tokensUpdated 2 days ago
    Business, Finance & HRAuto-check passed
  • Launch With Slurm

    mlc-ai/pith-train

    Reference for launching jobs inside a SLURM allocation via srun (single-node or multi-node).

    355 GitHub stars~1.4k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Launch Strategy

    alirezarezvani/claude-skills

    When the user wants to plan a product launch, feature announcement, or release strategy.

    28k GitHub stars~1.4k tokensUpdated 1 mo ago
    Marketing & SEOAuto-check passed

More from NVIDIA/skills

All 390 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.6k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.6k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.6k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.6k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.6k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.6k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemo Automodel Launcher Config

What does Nemo Automodel Launcher Config do?

Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution. Nemo Automodel Launcher Config is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.

How do I install Nemo Automodel Launcher Config in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-automodel-launcher-config -a claude-code`. Or copy the skill folder (skills/nemo-automodel-launcher-config in NVIDIA/skills) into .claude/skills/nemo-automodel-launcher-config in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Automodel Launcher Config in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-automodel-launcher-config -a codex`. Or copy the skill folder (skills/nemo-automodel-launcher-config in NVIDIA/skills) into .agents/skills/nemo-automodel-launcher-config in your project. Codex loads it when a task matches its description.

Can I use Nemo Automodel Launcher Config in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-automodel-launcher-config -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-automodel-launcher-config, .gemini/skills/nemo-automodel-launcher-config, .github/skills/nemo-automodel-launcher-config and .opencode/skills/nemo-automodel-launcher-config in your project.

What does Nemo Automodel Launcher Config need to run?

Going by SKILL.md and its folder, Nemo Automodel Launcher Config needs credentials named HF_TOKEN and WANDB_API_KEY. Our summary lists: Python 3; A credential in WANDB_API_KEY.

Does Nemo Automodel Launcher Config access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nemo Automodel Launcher Config safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Automodel Launcher Config use?

Nemo Automodel Launcher Config is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Automodel Launcher Config use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Automodel Launcher Config?

Skills that share tags, products or a category with Nemo Automodel Launcher Config: Configure Channel (openclaw/openclaw, 392k stars), Shipping and Launch Checklist (addyosmani/agent-skills, 105k stars), Slurm Job Script Generator (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars) and Technical Job Search (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Automodel Launcher Config?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.