Official agent skill

Nemo Automodel Recipe Development

by NVIDIA in NVIDIA/skills

Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.

OfficialApache-2.0Auto-check passed

Install Nemo Automodel Recipe Development

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-automodel-recipe-development -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-automodel-recipe-development --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-automodel-recipe-development .claude/skills/nemo-automodel-recipe-development && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-automodel-recipe-development
GitHub stars
3.5k
Token cost
~3k tokens
SKILL.md length
979 words
Files
5
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.

  • Works in 4 steps: Name the relevant recipe file or YAML… → List the builder functions or config… → Include a minimal YAML or command… → …
  • SKILL.md covers Instructions, Routing Boundary, Recipe Architecture and YAML Config Anatomy, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Nemo Automodel Recipe Development is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `BENCHMARK.md`, `evals/evals.json` and `skill-card.md`).

The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

Example prompts

  • “/nemo-automodel-recipe-development”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Name the relevant recipe file or YAML section.
  2. List the builder functions or config keys involved.
  3. Include a minimal YAML or command example when the question asks how to
  4. End with a local validation command or tiny CPU-compatible test.

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Automodel Recipe Development loads about 3k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 979 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 979 words, ~2,983 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-automodel-recipe-development/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
nemo-automodel-recipe-development
description
Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.
when_to_use
Creating or modifying training, SFT, or eval recipes, adding new YAML config fields, debugging recipe construction or trainer issues, or understanding the…
license
Apache-2.0
metadata.author
NVIDIA
metadata.tags
nemo-automodel, recipe-development

NeMo AutoModel Recipe Development

<!-- NVSkills signature refresh requested for AM-519. -->

Instructions

For recipe questions, answer with the smallest complete path to action:

  1. Name the relevant recipe file or YAML section.
  2. List the builder functions or config keys involved.
  3. Include a minimal YAML or command example when the question asks how to configure something.
  4. End with a local validation command or tiny CPU-compatible test.

For conceptual recipe questions, answer from this skill without inspecting the repository or loading other AutoModel skills unless the user asks you to edit files. Keep the response focused on recipe YAML, builders, CLI routing, tests, and local validation.

Use these compact answer patterns for common questions:

  • New finetuning recipe variant: start from the closest file under nemo_automodel/recipes/, update the model, dataset or dataloader, optimizer, loss, LR scheduler, step scheduler, and checkpoint builders, register a recipe alias only if adding a new recipe class, add example YAML under examples/, then add a tiny CPU-compatible unit test and run automodel <config.yaml>.
  • _target_ fields: describe _target_ as the fully qualified Python callable, explain that sibling keys become keyword arguments, show optimizer and dataset examples, and mention nested CLI overrides such as --optimizer.lr.
  • Validation and checkpointing: name step_scheduler.val_check_interval, step_scheduler.checkpoint_interval, validation_dataset, restore_from.path, and consolidated safetensors; include the minimal YAML snippet from this skill.

For validation and checkpointing, always name:

  • step_scheduler.val_check_interval for validation cadence.
  • step_scheduler.checkpoint_interval for save cadence.
  • validation_dataset as the validation dataloader source.
  • restore_from.path for resume.
  • Consolidated safetensors as the default checkpoint format for HF ecosystem compatibility.

Routing Boundary

Use this skill for recipe construction and execution-flow questions: YAML structure, _target_ callables, builder functions, validation datasets, checkpoint configuration, CLI route registration, and recipe-specific tests.

Do not use this skill for standalone distributed strategy selection, cluster launcher configuration, or model architecture onboarding unless the user is asking how those choices appear inside an AutoModel recipe YAML.

Recipe Architecture

Execution Flow
CLI (automodel config.yaml)
  -> app.py resolves the config's recipe target
    -> recipe script (e.g. train_ft.py) main(config_path)
      -> Recipe class .setup() builds all components
        -> .run_train_validation_loop() executes training
Recipe Class

Recipes inherit from BaseRecipe and implement two methods:

  • setup() -- builds model, optimizer, dataloader, loss, LR scheduler, step scheduler, and checkpoint config via builder functions.
  • run_train_validation_loop() -- executes the training and validation loop.
Builder Pattern

All components are constructed through dedicated builder functions:

  • build_model() -- instantiates the model from config
  • build_optimizer() -- creates optimizer (AdamW, etc.)
  • build_dataloader() -- sets up train and validation dataloaders
  • build_loss_module() -- creates the loss function
  • build_lr_scheduler() -- creates the learning rate scheduler
  • build_step_scheduler() -- creates the step scheduler controlling training progression
  • CheckpointingConfig -- configures checkpointing (built directly from the YAML checkpoint: block via RecipeConfig.checkpoint)
Infrastructure Application Order

Components are applied in this strict order after building:

  1. PEFT (LoRA, etc.)
  2. FP8 quantization
  3. QAT (quantization-aware training)
  4. Checkpoint load / restore
  5. Parameter freezing
  6. Sharding (FSDP2, Megatron-FSDP, DDP)
  7. Device placement
  8. torch.compile
  9. Context parallelism hooks

YAML Config Anatomy

A complete recipe config follows this structure:

yaml
step_scheduler:
  max_steps: 1000
  num_epochs: 1
  grad_accumulation_steps: 4
  val_check_interval: 100
  checkpoint_interval: 500
  log_interval: 10

dist_env:
  master_addr: localhost
  master_port: 29500

rng:
  seed: 42

model:
  _target_: nemo_automodel.NeMoAutoModelForCausalLM.from_pretrained
  pretrained_model_name_or_path: meta-llama/Llama-3.2-1B
  dtype: float32
  # additional model kwargs passed to the constructor

compile:
  enabled: false
  backend: inductor

clip_grad_norm:
  max_norm: 1.0

distributed:
  strategy: fsdp2       # fsdp2 | megatron_fsdp | ddp
  dp_size: auto
  tp_size: 1
  cp_size: 1

loss_fn:
  _target_: torch.nn.CrossEntropyLoss

dataset:
  _target_: nemo_automodel.datasets.squad.SquadDataset
  tokenizer_name_or_path: meta-llama/Llama-3.2-1B
  max_seq_length: 2048

validation_dataset:
  _target_: nemo_automodel.datasets.squad.SquadDataset
  split: validation

packed_sequence:
  enabled: false

dataloader:
  batch_size: 4
  num_workers: 4
  pin_memory: true

optimizer:
  _target_: torch.optim.AdamW
  lr: 2.0e-5
  weight_decay: 0.01

lr_scheduler:
  _target_: nemo_automodel.schedulers.CosineAnnealingWarmup
  warmup_steps: 50
  min_lr: 1.0e-6
Full-Parameter Training Precision

For new full-parameter training with torch.optim.Adam/AdamW, explicitly set model.dtype: float32 on NeMoAutoModel loaders for fp32 master weights and Adam moments. Configure compute precision separately (FSDP2: distributed.mp_policy).

PEFT, TE FusedAdam, and diffusion need separate precision choices; see the mixed-precision guide. Validate memory, training behavior, and checkpoint/resume when migrating existing configs.

The _target_ Pattern

The _target_ key specifies a fully qualified Python callable. All remaining keys in that section are passed as keyword arguments:

yaml
optimizer:
  _target_: torch.optim.AdamW   # callable
  lr: 2.0e-5                    # kwarg
  weight_decay: 0.01            # kwarg

This is equivalent to: torch.optim.AdamW(lr=2e-5, weight_decay=0.01).

CLI Overrides

Any config value can be overridden from the command line:

bash
automodel config.yaml \
  --optimizer.lr 1e-4 \
  --step_scheduler.max_steps 500 \
  --distributed.tp_size 2

Examples

Validation and checkpointing:

yaml
step_scheduler:
  val_check_interval: 100
  checkpoint_interval: 500

validation_dataset:
  _target_: nemo_automodel.datasets.squad.SquadDataset
  split: validation

restore_from:
  path: /checkpoints/step-500

Domain-Specific Notes

LLM
  • nemo_automodel/recipes/llm/train_ft.py handles both finetuning and pretraining. The distinction is in the config (dataset, learning rate, etc.).
  • nemo_automodel/recipes/llm/kd.py implements knowledge distillation with a teacher and student model.
  • nemo_automodel/recipes/llm/benchmark.py runs throughput and latency benchmarks.
Show full SKILL.md (395 more words)Show less
VLM
  • Uses NeMoAutoModelForImageTextToText instead of causal LM classes.
  • Config includes a processor section instead of a standalone tokenizer.
  • Recipe lives in nemo_automodel/recipes/vlm/finetune.py.
Diffusion
  • Uses NeMoAutoDiffusionPipeline.
  • Requires a parallel_scheme dict in config to define parallelism.
  • Only supports DDP and FSDP2 strategies (no Megatron-FSDP).
  • Recipe lives in nemo_automodel/recipes/diffusion/train.py.
Retrieval
  • Two encoder patterns:
    • Bi-encoder (nemo_automodel/recipes/retrieval/train_bi_encoder.py): separate query and document encoders, contrastive loss.
    • Cross-encoder (nemo_automodel/recipes/retrieval/train_cross_encoder.py): joint encoding, classification head.
  • Hard negative mining: nemo_automodel/recipes/retrieval/mine_hard_negatives.py.

Training Loop Details

The training loop follows this structure per epoch:

for epoch in range(num_epochs):
    for batch_idx in range(batches_per_epoch):
        # --- gradient accumulation inner loop ---
        for micro_batch in micro_batches:
            if pipeline_parallel:
                schedule.step(micro_batch)    # PP schedule
            else:
                loss = model(micro_batch)     # direct forward
                loss.backward()

        # --- optimizer step ---
        scale_grads_and_clip_grad_norm(model, max_norm)
        optimizer.step()
        lr_scheduler.step()
        optimizer.zero_grad()

        # --- logging ---
        MetricsSample(step, epoch, loss, grad_norm, lr, mem, tps, mfu)

        # --- validation (at configured intervals) ---
        if step % val_check_interval == 0:
            run_validation()

        # --- checkpoint (at configured intervals) ---
        if step % checkpoint_interval == 0:
            save_checkpoint()
StepScheduler

Controls all training progression: total epochs, total steps, gradient accumulation steps, validation interval, checkpoint interval, and logging interval.

Gradient Clipping

Applied via scale_grads_and_clip_grad_norm() after the backward pass and before the optimizer step. Controlled by clip_grad_norm.max_norm in config.

Context Parallelism

When cp_size > 1, batches are split across the context-parallel group using make_cp_batch_and_ctx(). This must happen before the forward pass.

MetricsSample

Each training step produces a MetricsSample with fields:

  • step -- global step count
  • epoch -- current epoch
  • loss -- training loss
  • grad_norm -- gradient norm after clipping
  • lr -- current learning rate
  • mem -- GPU memory usage
  • tps -- tokens per second
  • mfu -- model FLOPS utilization

Validation & Checkpointing

Validation
  • Runs at intervals defined by step_scheduler.val_check_interval.
  • Uses the validation dataloader built from validation_dataset config.
  • Model is set to eval mode; gradients are disabled.
Checkpointing
  • Default format: consolidated safetensors for easy deployment on HF ecosystem (always prefer this over DCP).
  • Checkpoint interval controlled by step_scheduler.checkpoint_interval.
  • Resume training via the restore_from config key pointing to a checkpoint directory.
yaml
restore_from:
  path: /checkpoints/step-500

Pitfalls

ProblemCauseFix
Silent config errorsTypo in _target_ valueThe class path must be a valid, importable Python callable. Double-check the module path and class name.
Training crashes at first stepglobal_batch_size not divisible by local_batch_size * dp_size * grad_accumulation_stepsEnsure the batch size math is consistent across all dimensions.
New recipe not accessible via CLIConfig is missing a resolvable recipe targetSet the config's recipe key to a discoverable recipe class name or a full dotted _target_ path.
Shape mismatch at forward passDataset collate function output does not match model input signatureVerify that the collate function returns tensors with the keys and shapes the model expects.
OOM during validationValidation batch size too large or gradients not disabledWrap validation in torch.no_grad() and consider a smaller validation batch size.
Checkpoint restore failsMismatched model architecture between checkpoint and configEnsure the model config matches the checkpoint exactly (layer count, hidden dim, vocab size).

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/nemo-automodel-recipe-development of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemo Automodel Recipe Development next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Automodel Recipe Development compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Automodel Recipe Development this skillNVIDIA/skills3.5k—~3kAutomated safety check: PassApache-2.0
Nemo Evaluator SDKOrchestra-Research/AI-Research-SKILLs13k2 repos~3.1kAutomated safety check: PassMIT
ML Training RecipesOrchestra-Research/AI-Research-SKILLs13k1 repos~2.8kAutomated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k1 repos~8.1kAutomated safety check: NotesMIT
Ito Trainingaffaan-m/ECC276k1 repos~1.5kAutomated safety check: PassMIT
LLM Evaluationdavila7/claude-code-templates32k12 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Nemo Evaluator SDK

    Orchestra-Research/AI-Research-SKILLs

    Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution.

    13k GitHub starsUsed in 2 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • ML Training Recipes

    Orchestra-Research/AI-Research-SKILLs

    PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM.

    13k GitHub starsUsed in 1 repo~2.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 1 repo~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Ito Training

    affaan-m/ECC

    Inspect the availability of ML training on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed training manifest.

    276k GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Train Pose

    ruvnet/RuView

    Train/evaluate WiFi pose models honestly — camera-supervised (MediaPipe + CSI) and camera-free (WiFlow), always checked against the mean-pose baseline before any PCK is quoted.

    97k GitHub stars~504 tokensUpdated today
    DevelopmentAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Nemo Automodel Recipe Development

What does Nemo Automodel Recipe Development do?

Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow. Nemo Automodel Recipe Development is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.

How do I install Nemo Automodel Recipe Development in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-automodel-recipe-development -a claude-code`. Or copy the skill folder (skills/nemo-automodel-recipe-development in NVIDIA/skills) into .claude/skills/nemo-automodel-recipe-development in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Automodel Recipe Development in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-automodel-recipe-development -a codex`. Or copy the skill folder (skills/nemo-automodel-recipe-development in NVIDIA/skills) into .agents/skills/nemo-automodel-recipe-development in your project. Codex loads it when a task matches its description.

Can I use Nemo Automodel Recipe Development in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-automodel-recipe-development -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-automodel-recipe-development, .gemini/skills/nemo-automodel-recipe-development, .github/skills/nemo-automodel-recipe-development and .opencode/skills/nemo-automodel-recipe-development in your project.

What does Nemo Automodel Recipe Development need to run?

SKILL.md names no scripts, command-line tools or credentials: Nemo Automodel Recipe Development is instructions for the agent only. Our summary lists: Python 3.

Does Nemo Automodel Recipe Development access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nemo Automodel Recipe Development safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Automodel Recipe Development use?

Nemo Automodel Recipe Development is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Automodel Recipe Development use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Automodel Recipe Development?

Skills that share tags, products or a category with Nemo Automodel Recipe Development: Nemo Evaluator SDK (Orchestra-Research/AI-Research-SKILLs, 13k stars), ML Training Recipes (Orchestra-Research/AI-Research-SKILLs, 13k stars), Arize Evaluator (github/awesome-copilot, 40k stars) and Ito Training (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Automodel Recipe Development?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.