Official agent skill

Kermt Pretrain Scratch

by NVIDIA in NVIDIA/skills

Pretrain a fresh KERMT model from scratch on a user-provided corpus.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Kermt Pretrain Scratch

skills CLI
$ npx skills add NVIDIA/skills --skill kermt-pretrain-scratch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills kermt-pretrain-scratch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/bionemo-kermt-pretrain-scratch .claude/skills/kermt-pretrain-scratch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
kermt-pretrain-scratch
GitHub stars
3.5k
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
955 words
Files
12 (incl. scripts)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Pretrain a fresh KERMT model from scratch on a user-provided corpus.

  • Works in 7 steps: Pre-flight: ensure container + system… → Compute run directory. → Validate the corpus (no ckpt to… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Skill and runtime paths, Hardware requirements, When to invoke and Inputs, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls python and git

What it does

Kermt Pretrain Scratch is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrainddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 14 other files, including scripts (for example `BENCHMARK.md`, `config/defaults_pretrain.json` and `evals/evals.json`). Compatibility notes: Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.

It sits in AI & LLM Engineering. It works with CUDA and NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/kermt-pretrain-scratch”

Requirements

  • Python 3
  • A Bash shell
  • Docker
  • Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Pre-flight: ensure container + system probe (same as
  2. Compute run directory.
  3. Validate the corpus (no ckpt to validate, so this is the only input
  4. Prepare the data — no vocab pass-through (we want fresh vocab from
  5. Estimate runtime + warn loudly. This is critical for pretrain-from-scratch
  6. Launch the runner detached.
  7. Report to the user. Always include all of the following — do not

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.

    From compatibility in the SKILL.md frontmatter.

Context cost

Kermt Pretrain Scratch loads about 2.4k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 955 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 955 words, ~2,377 tokens.

Download SKILL.mdSave it as .claude/skills/kermt-pretrain-scratch/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
kermt-pretrain-scratch
description
Pretrain a fresh KERMT model from scratch on a user-provided corpus. Builds a new vocabulary from the corpus, instantiates the model architecture from defaults, and launches pretrain_ddp.py inside the kermt container (detached for long runs). Unlike kermt-continue-pretrain, no starting checkpoint is loaded — the model is randomly initialized.
compatibility
Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron.
license
Apache-2.0
metadata.owner
evax@nvidia.com
metadata.classification
workflow-skill
metadata.risk_tier
skill

kermt-pretrain-scratch

Pretrain a brand-new KERMT model from scratch on a user-provided corpus. Useful when you want to retrain a model on a custom chemistry domain rather than extending one of the released checkpoints. Significantly more expensive than kermt-continue-pretrain — no warm start, so the loss curves need to descend from scratch over many epochs.

Skill and runtime paths

Set SKILL_DIR to the absolute path of this installed skill directory. Export KERMT_REPO as the absolute path to the KERMT checkout used for model execution. The bundled container helper mounts that checkout at /workspace and this skill at /skill (read-only). Commands inside the container use /skill/scripts/; defaults are bundled in config/.

Hardware requirements

Same as kermt-continue-pretrain:

  • GPUs: 1–N CUDA-capable. The runner auto-detects via torch.cuda.device_count(); --gpus 0,2 overrides. Single-GPU fallback: --batch_size 32 --save_interval 500. Multi-GPU keeps defaults (--batch_size 256 etc.). Note: --gpus N uses torch.cuda indexing, which can differ from nvidia-smi's display order on multi-GPU hosts (PCI bus vs. CUDA enumeration). To target a specific physical GPU, set CUDA_VISIBLE_DEVICES before invoking, or run python -c "import torch; print([torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())])" to confirm which device you're picking.

  • VRAM: the default --batch-size 256 is sized for A100-class hardware (80 GB VRAM). On smaller GPUs, downscale to avoid OOM:

    GPU classVRAMSuggested --batch-size
    L4, T4, V100 16 GB16–24 GB32–64
    A100 40 GB, L40, A4040–48 GB128
    A100 80 GB, H100, H20080 GB256 (default)

    These are rough starting points — pass --batch-size N to override.

  • Disk: tens of GB for shards + vocab + checkpoints, scaled by epochs.

  • Wall time: this is the big difference. Pretraining from scratch on an 11M-mol corpus at 100 epochs typically takes days even on a multi-GPU box. The skill prints an estimate before launching; confirm with the user.

When to invoke

  • User wants to train a new model on a custom corpus (e.g. domain-specific chemistry that the released ckpts don't cover).
  • User wants to reproduce a pretrain config end-to-end without depending on a released ckpt.

For continuing an existing released ckpt, use kermt-continue-pretrain. For adding a cMIM decoder to an encoder-only grover_base ckpt, use kermt-add-cmim-pretrain.

Inputs

Required:

  • --csv <path> — the pretrain corpus CSV with a smiles column. Single file by convention; multi-file corpora deferred. Use --val-csv for a separate validation set.
  • --pretrain-target-mode {vocab|cmim|hybrid} — which pretrain objective to use. No default — must be set explicitly so the user makes an informed choice:
    • vocab — original GROVER-style atom + bond vocab prediction (encoder-only output, lightweight).
    • cmim — contrastive + SMILES reconstruction objective. Requires building a SMILES vocab from the corpus.
    • hybrid — both vocab and contrastive objectives jointly (the state-of-the-art config from the KERMT manuscript).

Optional:

  • --val-csv <path> — separate validation CSV. Without it, prepare_data auto-splits the input by --val-frac 0.1 (random shuffle with --seed).
  • Training-hyperparameter overrides: --epochs N / --batch-size N / --init-lr F / --max-lr F / --final-lr F / --warmup-epochs F / --weight-decay F / --dropout F / --save-interval N / --seed N. Anything not given is filled from config/defaults_pretrain.json.
  • --vocab-loss-weight F (hybrid only) / --latent-dim N / --contrastive-temperature F (cmim and hybrid only).
  • --wandb-project NAME / --wandb-run-name NAME — optional Weights & Biases logging. When --wandb-project is set, rank 0 logs train/val losses; the run name is honored only alongside a project. Off by default.
  • --gpus 0,2 — restrict to a GPU subset.
Show full SKILL.md (480 more words)Show less

Workflow

Let $KERMT_REPO be the path to your kermt repo checkout.

  1. Pre-flight: ensure container + system probe (same as kermt-continue-pretrain step 1). Refuse to proceed if check_system reports gaps.

  2. Compute run directory.

    RUN_DIR=$KERMT_REPO/runs/pretrain-scratch_$(date -u +%Y-%m-%dT%H-%M-%SZ)
  3. Validate the corpus (no ckpt to validate, so this is the only input check):

    "$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> -- \
        "python /skill/scripts/check_data.py --mode pretrain --csv /data/<basename>"

    Abort on ok: false.

  4. Prepare the data — no vocab pass-through (we want fresh vocab from corpus):

    "$SKILL_DIR/scripts/kermt_container.sh" run --data <user-csv> --run-dir $RUN_DIR -- \
        "python /skill/scripts/prepare_data.py --mode pretrain \\
             --csv /data/<basename> --out /runs/data \\
             [--val-csv /data/<val-basename>] [--val-frac 0.1] [--seed 0]"

    Outputs land at $RUN_DIR/data/prepare_data.json with vocab_source: "built_fresh".

  5. Estimate runtime + warn loudly. This is critical for pretrain-from-scratch:

    • "Pretraining from scratch is days-scale even on multi-GPU; the released KERMT checkpoints were each trained on millions of molecules for hundreds of GPU-hours. If you mainly want to leverage existing knowledge for a downstream task, consider kermt-continue-pretrain from a released ckpt instead, which converges in hours instead of days."
    • Show the corpus size × epochs × GPU count → estimated wall time.
    • Ask for explicit confirmation unless --yes was given.
  6. Launch the runner detached.

    "$SKILL_DIR/scripts/kermt_container.sh" run_detached \\
        --name kermt-pretrain-scratch-<ts> \\
        --run-dir $RUN_DIR -- \\
        "python /skill/scripts/run_pretrain_local.py \\
             --from-scratch --pretrain-target-mode <vocab|cmim|hybrid> \\
             --prepare-manifest /runs/data/prepare_data.json \\
             --out /runs \\
             [--epochs N --batch-size N ...]"

    Note: NO --ckpt flag (the runner refuses if both --from-scratch and --ckpt are given). The runner uses the arch group from config/defaults_pretrain.json to size the model.

  7. Report to the user. Always include all of the following — do not omit the TensorBoard line under output-length pressure:

    • Container name + id
    • $RUN_DIR/run.json (the manifest with workflow: pretrain-scratch, from_scratch: true, vocab_check: null, arch from defaults, full cmd_replay)
    • Log file: $RUN_DIR/logs/pretrain_ddp.log
    • TensorBoard: $RUN_DIR/logs/tb (open with tensorboard --logdir $RUN_DIR/logs/tb)
    • Suggest kermt-monitor <RUN_DIR> for progress.

Hard rules

  • Never accept a --ckpt flag. From-scratch is exclusive with input ckpt — the runner enforces this; the skill should too.
  • Never silently default --pretrain-target-mode. This is a significant architectural choice (vocab = lightweight, hybrid = SOTA). Prompt the user if not given on the CLI.
  • Strong warning before launching. From-scratch pretrain is the most expensive workflow. The user needs to know what they're committing to.

Common errors

  • --pretrain-target-mode is required when --from-scratch is set → user forgot the mode flag. Prompt.
  • --from-scratch is incompatible with --ckpt → user provided both; ask which one they meant.
  • defaults_pretrain.json has no arch group → repo state issue (should never happen on a fresh clone); points the user at running kermt-setup again.

What's in the manifest after a from-scratch run

Same reproducibility fields as continue-pretrain (repo.commit, kermt_image, cmd_replay, args_applied), plus:

  • workflow: "pretrain-scratch"
  • from_scratch: true
  • inputs.ckpt: null
  • ckpt_symlink: null
  • vocab_check: null (not verified — vocab built from corpus is authoritative for from-scratch)
  • arch: the values pulled from config/defaults_pretrain.json's arch group (with any future CLI overrides applied).

Replayability

Same as continue-pretrain: cmd_replay is a copy-pasteable command. If ok_to_replay: false, the kermt repo working tree was dirty at launch time — check repo.commit and git checkout it first.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (scripts) in skills/bionemo-kermt-pretrain-scratch of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • config/defaults_pretrain.json
  • evals/evals.json
  • scripts/_utils.py
  • scripts/check_checkpoint.py
  • scripts/check_data.py
  • scripts/kermt_container.sh
  • scripts/prepare_data.py
  • scripts/run_pretrain_local.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 0e0d506

Used in 2 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Kermt Pretrain Scratch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Kermt Pretrain Scratch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Kermt Pretrain Scratch this skillNVIDIA/skills3.5k1 repos~2.4kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNone
Optimize OpCVCUDA/CV-CUDA2.7k—~834Automated safety check: PassCustom licence
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent143—~5.2kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 9 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    900 GitHub stars~2.8k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Optimize Op

    CVCUDA/CV-CUDA

    Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

    2.7k GitHub stars~834 tokensUpdated 20 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    143 GitHub stars~5.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Llama Cpp

    Orchestra-Research/AI-Research-SKILLs

    Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

    13k GitHub starsUsed in 4 repos~1.5k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Kermt Pretrain Scratch

What does Kermt Pretrain Scratch do?

Pretrain a fresh KERMT model from scratch on a user-provided corpus. Kermt Pretrain Scratch is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Pretrain a fresh KERMT model from scratch on a user-provided corpus.

When should I use Kermt Pretrain Scratch?

Kermt Pretrain Scratch fits situations like: AI & LLM Engineering work in your project.

How do I install Kermt Pretrain Scratch in Claude Code?

Run `npx skills add NVIDIA/skills --skill kermt-pretrain-scratch -a claude-code`. Or copy the skill folder (skills/bionemo-kermt-pretrain-scratch in NVIDIA/skills) into .claude/skills/kermt-pretrain-scratch in your project. Claude Code loads it when a task matches its description.

How do I install Kermt Pretrain Scratch in Codex?

Run `npx skills add NVIDIA/skills --skill kermt-pretrain-scratch -a codex`. Or copy the skill folder (skills/bionemo-kermt-pretrain-scratch in NVIDIA/skills) into .agents/skills/kermt-pretrain-scratch in your project. Codex loads it when a task matches its description.

Can I use Kermt Pretrain Scratch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill kermt-pretrain-scratch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kermt-pretrain-scratch, .gemini/skills/kermt-pretrain-scratch, .github/skills/kermt-pretrain-scratch and .opencode/skills/kermt-pretrain-scratch in your project.

What does Kermt Pretrain Scratch need to run?

Going by SKILL.md and its folder, Kermt Pretrain Scratch needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python and git). Our summary lists: Python 3; A Bash shell; Docker. Compatibility (from SKILL.md): Requires docker, nvidia-container-toolkit, and a CUDA-capable NVIDIA GPU. Designed for Claude Code, Codex, and Nemotron..

Does Kermt Pretrain Scratch access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Kermt Pretrain Scratch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Kermt Pretrain Scratch use?

Kermt Pretrain Scratch is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Kermt Pretrain Scratch use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Kermt Pretrain Scratch?

Skills that share tags, products or a category with Kermt Pretrain Scratch: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 900 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Kermt Pretrain Scratch?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.