Official agent skill

Megatron GB200 One-Node Test Onboarding

by NVIDIA in NVIDIA/Megatron-LM

Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.

OfficialApache-2.0Auto-check passedTesting & QA

Install Megatron GB200 One-Node Test Onboarding

skills CLI
$ npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/Megatron-LM mcore-onboard-gb200-1node-tests --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mcore-onboard-gb200-1node-tests .claude/skills/mcore-onboard-gb200-1node-tests && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mcore-onboard-gb200-1node-tests
GitHub stars
18k
Token cost
~1.3k tokens
SKILL.md length
486 words
Files
5
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.

  • Works in 6 steps: Find candidate tests → Read each model config → Classify: trivial copy vs. needs… → …
  • Adding one-node GitHub MR coverage for GB200 functional tests
  • SKILL.md covers Background, Workflow, Quick parallelism reference and Checklist
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Each GB200 node has 4 GPUs, so a two-node test uses 8 and its one-node variant uses 4. The skill scans the products block of gpt.yaml and moe.yaml for tests scoped mr or mr-slim, skips ones already covered in the 1node recipe files and anything nightly, weekly or mr-broken, then reads each model_config.yaml for the tensor, pipeline and expert parallel sizes.

Each candidate is classified using world size as TP times PP times DP. If TP times PP is at most 4, the config is copied unchanged and DP halves on its own; if it equals 8, pipeline parallelism is halved; expert parallelism above 4 is reduced to 4; and ETP tests are re-checked afterwards. Global batch size is left alone so gradient accumulation absorbs the lower DP. New configs go in test case folders with a _1node suffix, and recipes live under tests/test_utils/recipes/gb200. The folder also carries a benchmark file, a skill card and evals.

When your agent uses it

  • Adding one-node GitHub MR coverage for GB200 functional tests
  • Deciding whether a two-node test config can be copied or needs smaller parallelism
  • Creating a gpt-1node.yaml recipe when none exists

Example prompts

  • “Find the mr-scoped two-node GB200 GPT tests that still lack a one-node variant.”
  • “Create the _1node model_config for this MoE test; it uses expert parallel size 8.”
  • “Check whether tensor parallel 4 with pipeline parallel 2 fits on a single GB200 node.”

Requirements

  • A checkout of the Megatron-LM repository with the GB200 test recipes

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Find candidate tests
  2. Read each model config
  3. Classify: trivial copy vs. needs adaptation
  4. Create _1node model config directories
  5. Create or update recipe files
  6. Add products entries

What it can do on your machine

Read from SKILL.md and the folder at commit d5fbb65. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Megatron GB200 One-Node Test Onboarding loads about 1.3k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 486 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~30
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/Megatron-LM at commit d5fbb65, republished under its Apache-2.0 licence (© NVIDIA). 486 words, ~1,339 tokens.

Download SKILL.mdSave it as .claude/skills/mcore-onboard-gb200-1node-tests/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
mcore-onboard-gb200-1node-tests
description
Onboard 1-node GitHub MR functional tests for GB200 from existing mr-scoped 2-node tests.
license
Apache-2.0
when_to_use
Adding GB200 github-mr tests; creating single-node variants of existing tests; expanding CI coverage for GB200; 'add GB200 MR tests', 'onboard GB200 1-node'…
user_invocable
true
argument
[model-yaml] # optional: gpt, moe, or both (default: both)
metadata.author
Oliver Koenig <okoenig@nvidia.com>

Onboard GB200 1-Node GitHub MR Tests

Create 1-node (mr-github) variants of existing 2-node (mr-scoped) GB200 functional tests. Each GB200 node has 4 GPUs. A 2-node test uses 8 GPUs total; the 1-node variant uses 4.


Background

GB200 functional tests live in tests/test_utils/recipes/gb200/:

Recipe fileNotes
gpt.yamlGPT dense tests, nodes: 2, gpus: 4 (8 total)
moe.yamlMoE tests, nodes: 2, gpus: 4 (8 total)
moe-1node.yamlExisting 1-node MoE tests, nodes: 1, gpus: 4 (4 total)
gpt-1node.yaml1-node GPT tests (create if not present)

Model configs live at: tests/functional_tests/test_cases/{model}/{test_case}/model_config.yaml

1-node test cases use the _1node suffix: tests/functional_tests/test_cases/{model}/{test_case}_1node/model_config.yaml


Workflow

Step 1 — Find candidate tests

Scan the products: block in gpt.yaml and moe.yaml for entries with scope: [mr, ...] or scope: [mr-slim, ...]. These are the 2-node tests that need 1-node mr-github counterparts.

Ignore tests already covered in *-1node.yaml files, and ignore nightly, weekly, mr-broken scopes.

Step 2 — Read each model config

For each candidate, read its model_config.yaml and extract the key parallelism arguments:

--tensor-model-parallel-size   (TP)
--pipeline-model-parallel-size (PP)
--expert-model-parallel-size   (EP)
--expert-tensor-parallel-size  (ETP)
--context-parallel-size        (CP)
--global-batch-size
--micro-batch-size
Step 3 — Classify: trivial copy vs. needs adaptation

The world size formula is: world_size = TP × PP × DP where DP ≥ EP.

Going from 8 GPUs → 4 GPUs:

ConditionAction
TP × PP ≤ 4Trivial copy. Config unchanged; DP is halved automatically.
TP × PP = 8 (e.g. tp4 pp2)Reduce PP. Set PP = PP / 2 (e.g. pp2→1). Verify TP × PP_new ≤ 4.
EP > 4 (e.g. ep8 with tp1 pp1)Reduce EP. Set EP = 4. Experts stay at num-experts (each EP rank holds more experts).
EP > 4 and TP × PP > 4Reduce both PP and EP as above.
ETP test (ep × etp ≤ TP × DP)Check EP × ETP ≤ TP × DP_new after PP reduction. Usually satisfied when pp→1.

Do not change GBS — let gradient accumulation absorb the reduced DP.

Step 4 — Create _1node model config directories
bash
# Trivial copy
mkdir -p tests/functional_tests/test_cases/{model}/{test_case}_1node
cp tests/functional_tests/test_cases/{model}/{test_case}/model_config.yaml \
   tests/functional_tests/test_cases/{model}/{test_case}_1node/model_config.yaml

# Then apply any parallelism changes (EP or PP) with Edit tool
Show full SKILL.md (199 more words)Show less
Step 5 — Create or update recipe files

For GPT tests — create tests/test_utils/recipes/gb200/gpt-1node.yaml (if absent) by cloning gpt.yaml's spec block with nodes: 1. Use this template for the spec:

yaml
type: basic
format_version: 1
maintainers: [mcore]
loggers: [stdout]
spec:
  name: "{test_case}_{environment}_{platforms}"
  model: gpt          # or moe
  build: mcore-pyt-{environment}
  nodes: 1
  gpus: 4
  n_repeat: 5
  platforms: dgx_gb200
  script_setup: |    # copy verbatim from gpt.yaml / moe.yaml
    ...
  script: |-         # copy verbatim from gpt.yaml / moe.yaml
    ...

For MoE tests — append entries to the existing moe-1node.yaml.

Step 6 — Add products entries

Scope convention:

  • 1–2 most representative tests per recipe: scope: [mr-github, mr-github-slim]
  • All other tests: scope: [mr-github]
yaml
products:
  - test_case: [<test_case>_1node]
    products:
      - environment: [dev]
        scope: [mr-github, mr-github-slim]   # or [mr-github]
        platforms: [dgx_gb200]

Quick parallelism reference

Original (8 GPUs)1-node config (4 GPUs)Notes
tp1 pp1 ep1 → dp8tp1 pp1 ep1 → dp4trivial
tp2 pp1 ep1 → dp4tp2 pp1 ep1 → dp2trivial
tp1 pp2 ep1 → dp4tp1 pp2 ep1 → dp2trivial
tp4 pp1 ep1 → dp2tp4 pp1 ep1 → dp1trivial
tp1 pp4 ep1 → dp2tp1 pp4 ep1 → dp1trivial
tp1 pp1 ep8 → dp8tp1 pp1 ep4 → dp4ep 8→4
tp4 pp2 ep2 etp2 → dp1tp4 pp1 ep2 etp2 → dp1pp 2→1

Checklist

  • Identified all mr-scoped tests in gpt.yaml and moe.yaml not yet in *-1node.yaml
  • Read model config for each candidate
  • Classified trivial vs. adaptation needed
  • Created _1node/model_config.yaml for each test
  • Applied EP or PP reductions where needed
  • Created/updated recipe YAML with nodes: 1, gpus: 4
  • Assigned mr-github scope (+ mr-github-slim for 1–2 representative tests per recipe)
  • Verified no mr-github-slim overload (slim suite should stay small)

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/mcore-onboard-gb200-1node-tests of NVIDIA/Megatron-LM.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit d5fbb65

Compare with similar skills

Megatron GB200 One-Node Test Onboarding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Megatron GB200 One-Node Test Onboarding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Megatron GB200 One-Node Test Onboarding this skillNVIDIA/Megatron-LM18k—~1.3kAutomated safety check: PassApache-2.0
Explore Feature E2E Testcomet-ml/opik22k—~3.4kAutomated safety check: PassApache-2.0
Kane CLI Browser TestingLambdaTest/kane-cli247—~8.4kAutomated safety check: PassApache-2.0
Nemoclaw Maintainer Analyze CI PerformanceNVIDIA/NemoClaw23k—~644Automated safety check: PassApache-2.0
Endgamemicrosoft/copilot-for-eclipse126—~1.4kAutomated safety check: PassMIT
Implementnvuillam/github-dependents-info162—~999Automated safety check: PassMIT

Similar skills

  • Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.

    22k GitHub stars~3.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Kane CLI Browser Testing

    LambdaTest/kane-cli

    Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.

    247 GitHub stars~8.4k tokensUpdated 3 days ago
    Testing & QAAuto-check passed
  • Analyze retained NemoClaw CI timings for slow CLI tests, runner queues, or base-image publication.

    23k GitHub stars~644 tokensUpdated today
    Testing & QAAuto-check passed
  • Endgame

    microsoft/copilot-for-eclipse

    Official

    Orchestrate endgame verification for a GitHub milestone issue.

    126 GitHub stars~1.4k tokensUpdated 8 days ago
    Testing & QAAuto-check passed
  • Implement

    nvuillam/github-dependents-info

    Phase 3 of the SDLC pipeline (also usable standalone). An agent skill from nvuillam/github-dependents-info.

    162 GitHub stars~999 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Official

    Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug.

    23k GitHub stars~3k tokensUpdated today
    Testing & QAAuto-check passed

More from NVIDIA/Megatron-LM

All 14 skills in this repo
  • Megatron Core Testing Guide

    NVIDIA/Megatron-LM

    Official

    Guide to the Megatron-LM test system: layout, recipe YAML, running and adding unit and functional tests, golden values, marker filters and CI parity.

    18k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Official

    Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.

    18k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Official

    Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.

    18k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Megatron-LM CI/CD Guide

    NVIDIA/Megatron-LM

    Official

    Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.

    18k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Official

    Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.

    18k GitHub stars~1.6k tokensUpdated today
    Auto-check passed

Categories

Questions about Megatron GB200 One-Node Test Onboarding

What does Megatron GB200 One-Node Test Onboarding do?

Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs. Each GB200 node has 4 GPUs, so a two-node test uses 8 and its one-node variant uses 4.yaml for the tensor, pipeline and expert parallel sizes.

When should I use Megatron GB200 One-Node Test Onboarding?

Megatron GB200 One-Node Test Onboarding fits situations like: adding one-node GitHub MR coverage for GB200 functional tests; deciding whether a two-node test config can be copied or needs smaller parallelism; creating a gpt-1node.yaml recipe when none exists.

How do I install Megatron GB200 One-Node Test Onboarding in Claude Code?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a claude-code`. Or copy the skill folder (skills/mcore-onboard-gb200-1node-tests in NVIDIA/Megatron-LM) into .claude/skills/mcore-onboard-gb200-1node-tests in your project. Claude Code loads it when a task matches its description.

How do I install Megatron GB200 One-Node Test Onboarding in Codex?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a codex`. Or copy the skill folder (skills/mcore-onboard-gb200-1node-tests in NVIDIA/Megatron-LM) into .agents/skills/mcore-onboard-gb200-1node-tests in your project. Codex loads it when a task matches its description.

Can I use Megatron GB200 One-Node Test Onboarding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mcore-onboard-gb200-1node-tests, .gemini/skills/mcore-onboard-gb200-1node-tests, .github/skills/mcore-onboard-gb200-1node-tests and .opencode/skills/mcore-onboard-gb200-1node-tests in your project.

What does Megatron GB200 One-Node Test Onboarding need to run?

SKILL.md names no scripts, command-line tools or credentials: Megatron GB200 One-Node Test Onboarding is instructions for the agent only. Our summary lists: A checkout of the Megatron-LM repository with the GB200 test recipes.

Does Megatron GB200 One-Node Test Onboarding access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Megatron GB200 One-Node Test Onboarding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Megatron GB200 One-Node Test Onboarding use?

Megatron GB200 One-Node Test Onboarding is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Megatron GB200 One-Node Test Onboarding use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Megatron GB200 One-Node Test Onboarding?

Skills that share tags, products or a category with Megatron GB200 One-Node Test Onboarding: Explore Feature E2E Test (comet-ml/opik, 22k stars), Kane CLI Browser Testing (LambdaTest/kane-cli, 247 stars), Nemoclaw Maintainer Analyze CI Performance (NVIDIA/NemoClaw, 23k stars) and Endgame (microsoft/copilot-for-eclipse, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Megatron GB200 One-Node Test Onboarding?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/Megatron-LM, which has 18,083 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.

Source: NVIDIA/Megatron-LM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.