Official agent skill

Megatron Core Testing Guide

by NVIDIA in NVIDIA/Megatron-LM

Guide to the Megatron-LM test system: layout, recipe YAML, running and adding unit and functional tests, golden values, marker filters and CI parity.

OfficialApache-2.0Auto-check passedTesting & QA

Install Megatron Core Testing Guide

skills CLI
$ npx skills add NVIDIA/Megatron-LM --skill mcore-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/Megatron-LM mcore-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mcore-testing .claude/skills/mcore-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mcore-testing
GitHub stars
18k
Token cost
~2.1k tokens
SKILL.md length
678 words
Files
5
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide to the Megatron-LM test system: layout, recipe YAML, running and adding unit and functional tests, golden values, marker filters and CI parity.

  • Works in 6 steps: Create tests/unit_tests//test_.py. → Use fixtures from… → Apply markers as needed → …
  • Adding a unit or functional test to Megatron-LM
  • SKILL.md covers Answer-First Testing Facts, Test Layout, How Tests Execute and Recipe YAML Structure, plus 4 more sections
  • Calls uv and python

What it does

The guide opens with answer-first facts on disabling tests without deleting them. In functional recipe YAML you suffix the scope with -broken, for example turning mr-github into mr-github-broken, while unit tests use pytest markers: flaky_in_dev skips in the default dev environment and flaky skips in LTS. The test case or recipe entry itself is kept so it stays discoverable and easy to re-enable.

It then explains how tests execute. A GitHub Actions runner calls launch_nemo_run_workload.py, which uses nemo-run to start a Docker container with the repo and training data mounted. Unit tests run through torch.distributed.run on one node with 8 GPUs and log per rank, while functional tests are driven by run_ci_test.sh and only rank 0 runs the pytest validation. Known transient failures such as NCCL timeouts, ECC errors, segfaults and HuggingFace connectivity are retried up to 3 times.

Recipes live in tests/test_utils/recipes/ and are parsed by recipe_parser.py, which expands a products block into individual workload specs with runtime placeholders. Because every unit test initializes a torch.distributed group, local runs need GPU access. A benchmark file and an evals file ship with the skill.

When your agent uses it

  • Adding a unit or functional test to Megatron-LM
  • Temporarily disabling a flaky test without deleting it
  • Reproducing a CI test failure locally and checking CI parity

Example prompts

  • “Disable the flaky functional test for the GPT recipe without deleting its entry.”
  • “How do I run the distributed optimizer unit tests locally on one node?”
  • “Add a new functional test case to the recipe YAML and explain how CI picks it up.”

Requirements

  • A Megatron-LM checkout
  • GPU access for unit tests
  • Docker for CI-style runs

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Create tests/unit_tests//test_.py.
  2. Use fixtures from tests/unit_tests/conftest.py.
  3. Apply markers as needed
  4. Verify locally (see Running Unit Tests Locally above).
  5. If the test needs a dedicated CI bucket, add an entry to
  6. If the change adds or modifies a GPU kernel (Triton, jit_fuser /

What it can do on your machine

Read from SKILL.md and the folder at commit 486a126. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Megatron Core Testing Guide loads about 2.1k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 678 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/Megatron-LM at commit 486a126, republished under its Apache-2.0 licence (© NVIDIA). 678 words, ~2,131 tokens.

Download SKILL.mdSave it as .claude/skills/mcore-testing/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
mcore-testing
description
Test system for Megatron-LM. Covers test layout, recipe YAML structure, adding and running unit and functional tests, golden values, marker filters, and CI parity.
license
Apache-2.0
when_to_use
Adding or running a unit or functional test; understanding the test layout; writing a recipe YAML; downloading or updating golden values; reproducing a test…
metadata.author
Oliver Koenig <okoenig@nvidia.com>

Testing Guide


Answer-First Testing Facts

For questions about disabling tests without deleting them:

  • Functional recipe entries stay in YAML; disable by suffixing scope with -broken, for example scope: [mr-github] -> scope: [mr-github-broken].
  • Unit-test skips use pytest markers instead: @pytest.mark.flaky_in_dev skips in the default dev environment, and @pytest.mark.flaky skips in LTS.
  • Do not delete the test case or recipe entry when the goal is discoverability and easy re-enable.

Test Layout

text
tests/
├── unit_tests/          # pytest, 1 node × 8 GPUs, torch.distributed runner
├── functional_tests/    # end-to-end shell + training scripts
│   └── test_cases/
│       └── {model}/{test_case}/
│           ├── model_config.yaml          # training args
│           └── golden_values_{env}_{platform}.json
└── test_utils/
    ├── recipes/
    │   ├── h100/        # YAML recipes for H100 jobs
    │   └── gb200/       # YAML recipes for GB200 jobs
    └── python_scripts/  # helpers (recipe_parser, golden-value download, …)

How Tests Execute

The GitHub Actions runner invokes launch_nemo_run_workload.py, which uses nemo-run to launch a DockerExecutor container. The repo is bind-mounted at /opt/megatron-lm; training data is mounted at /mnt/artifacts.

Unit tests are dispatched through torch.distributed.run:

  • Ranks 0 and 3 are tee-d to stdout; all other ranks write only to log files.
  • Per-rank log files land at {assets_dir}/logs/1/ and are uploaded as a GitHub artifact after the run.

Functional tests are driven by tests/functional_tests/shell_test_utils/run_ci_test.sh. Only rank 0 runs the pytest validation step; training output from all ranks is uploaded as an artifact.

Flaky-failure auto-retry: launch_nemo_run_workload.py retries up to 3 times for known transient patterns (NCCL timeout, ECC error, segfault, HuggingFace connectivity, …) before declaring a genuine failure.


Recipe YAML Structure

Recipes live in tests/test_utils/recipes/ and are parsed by tests/test_utils/python_scripts/recipe_parser.py. Each file expands a cartesian products block into individual workload specs:

yaml
type: basic
format_version: 1
maintainers: [mcore]
loggers: [stdout]
spec:
  name: "{test_case}_{environment}_{platforms}"
  model: gpt              # maps to tests/functional_tests/test_cases/{model}/
  build: mcore-pyt-{environment}
  nodes: 1
  gpus: 8
  n_repeat: 5
  platforms: dgx_h100
  time_limit: 1800
  script_setup: |
    ...
  script: |-
    bash tests/functional_tests/shell_test_utils/run_ci_test.sh ...
products:
  - test_case: [my_test]
    products:
      - environment: [dev, lts]
        scope: [mr-github]
        platforms: [dgx_h100]

Key runtime placeholders: {assets_dir}, {artifacts_dir}, {test_case}, {environment}, {platforms}, {n_repeat}.

Disabling a Test Without Deleting It

To temporarily disable a test case in a recipe YAML, suffix its scope value with -broken — do not delete the entry:

yaml
# before (test runs in CI)
scope: [mr-github]

# after (test is skipped; entry preserved for easy re-enable)
scope: [mr-github-broken]

Running Unit Tests Locally

All unit tests initialize a torch.distributed group, so every invocation requires GPU access and must go through torch.distributed.run:

bash
# Full suite
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
  tests/unit_tests

# Single file
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
  tests/unit_tests/models/test_gpt_model.py

# Single test
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
  tests/unit_tests/models/test_gpt_model.py::TestGPTModel::test_constructor

# Filter by name substring
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
  tests/unit_tests -k optimizer
Marker filters
bash
# Exclude flaky tests during development
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
  tests/unit_tests -m "not flaky and not flaky_in_dev"

# Include experimental tests
uv run python -m torch.distributed.run --nproc-per-node 8 -m pytest -q \
  tests/unit_tests --experimental
CI parity

Use tests/unit_tests/run_ci_test.sh to reproduce a CI bucket failure exactly. For ad-hoc runs, prefer the direct torch.distributed.run invocations above.

Gotchas
  • pyproject.toml sets addopts = --durations=15 -s -rA — stdout is not captured (-s), so ranks interleave during multi-rank runs. Override with --capture=fd when debugging a specific rank.
  • tests/unit_tests/conftest.py looks for test data under /opt/data and attempts a download if missing. Supply it manually or skip data-dependent tests when running outside the canonical container.

Adding a Unit Test

  1. Create tests/unit_tests/<category>/test_<name>.py.
  2. Use fixtures from tests/unit_tests/conftest.py.
  3. Apply markers as needed:
    • @pytest.mark.internal — skipped on legacy tag
    • @pytest.mark.flaky_in_dev — skipped in dev environment (CI default; use this to disable a flaky test without blocking the standard pipeline)
    • @pytest.mark.flaky — skipped in lts environment
    • @pytest.mark.experimental — latest tag only
  4. Verify locally (see Running Unit Tests Locally above).
  5. If the test needs a dedicated CI bucket, add an entry to tests/test_utils/recipes/h100/unit-tests.yaml.
  6. If the change adds or modifies a GPU kernel (Triton, jit_fuser / torch.compile, CUDA extension, TE or external-library dispatch, or a scatter/index accumulation), add or update its bit-exact replay test under tests/unit_tests/determinism/kernels/ and register it in tests/unit_tests/determinism/kernels/manifest.py. The linting CI job (tools/check_kernel_determinism_coverage.py) fails kernel PRs without this; the determinism-exempt label overrides it for non-numeric edits. See docs/developer/determinism/testing.md.

Show full SKILL.md (217 more words)Show less

Adding a Functional / Integration Test

  1. Create tests/functional_tests/test_cases/<model>/<test_name>/.

  2. Write model_config.yaml with MODEL_ARGS, ENV_VARS, and TEST_TYPE.

  3. Add a YAML recipe under tests/test_utils/recipes/h100/ (and gb200/ if needed). Required fields: scope, environment, platform, n_repeat, time_limit.

  4. Push the PR, add the label "Run functional tests" to trigger a full run.

  5. After a successful run, download golden values:

    bash
    python tests/test_utils/python_scripts/download_golden_values.py \
      --source github --pipeline-id <run-id>
  6. Commit the downloaded golden values.

Golden values keep the full float32 precision of the TensorBoard scalars, record it as "value_precision": "full", and deterministic test cases are compared bit-exactly against them. Never round or hand-edit them: tools/check_golden_values.py (run by the linting CI job on changed golden files) rejects golden files of deterministically compared cases whose metrics lack the full marker. Legacy files (no marker) are still compared at five decimals until they are regenerated from a CI run.


Golden comparisons report absolute and relative errors; see diagnostics.

Common Pitfalls

ProblemCauseFix
Test passes locally but fails in CIDifferent environment or data pathCheck DATA_PATH, DATA_CACHE_PATH, and the environment tag (dev vs lts)
Golden value mismatch after a code changeNumerical regressionDownload new golden values via download_golden_values.py after a clean run
cicd-integration-tests-gb200 not triggeredGB200 jobs require maintainer statusAsk a maintainer to trigger, or add the Run functional tests label

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/mcore-testing of NVIDIA/Megatron-LM.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 486a126

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in NVIDIA/Megatron-LM, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Megatron Core Testing Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Megatron Core Testing Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Megatron Core Testing Guide this skillNVIDIA/Megatron-LM18k—~2.1kAutomated safety check: PassApache-2.0
Simple Modern Uvjlevy/simple-modern-uv301—~1.9kAutomated safety check: PassMIT
Robotics Testingarpitg1304/robotics-agent-skills368—~4.7kAutomated safety check: PassApache-2.0
Sitl TestingArduPilot/MethodicConfigurator163—~1.7kAutomated safety check: PassGPL-3.0
Gating Deid Leakagemaziyarpanahi/openmed5.5k—~1.7kAutomated safety check: PassApache-2.0
Code PatternsAedelon/claude-code-blueprint120—~1.2kAutomated safety check: PassCustom licence

Similar skills

  • Simple Modern Uv

    jlevy/simple-modern-uv

    Start, selectively modernize, fully migrate, or update Python projects using simple-modern-uv practices: uv, ruff, BasedPyright, pytest, GitHub Actions CI, and tag-driven PyPI publishing.

    301 GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Robotics Testing

    arpitg1304/robotics-agent-skills

    Testing strategies, patterns, and tools for robotics software.

    368 GitHub stars~4.7k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Sitl Testing

    ArduPilot/MethodicConfigurator

    Set up and run SITL integration tests for backendflightcontroller.py.

    163 GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Gating Deid Leakage

    maziyarpanahi/openmed

    Add a CI gate that fails the build when an OpenMed de-identification model's recall on a held-out PHI set drops below threshold or any critical identifier leaks.

    5.5k GitHub stars~1.7k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Code Patterns

    Aedelon/claude-code-blueprint

    Reference patterns for REST APIs, pytest/vitest testing, Docker multi-stage builds, GitHub Actions CI/CD, PostgreSQL, TypeScript generics, Python async, and React Server Components.

    120 GitHub stars~1.2k tokensUpdated 7 mo ago
    DevOps & CloudAuto-check passed
  • Django CI Test Optimization

    hashgraph-online/awesome-codex-plugins

    Optimize Django and pytest-django test execution in CI with cache configuration, slow-test splitting, database reuse strategy, parallel workers, pytest-xdist, CircleCI/GitHub Actions/Jenkins/Travis…

    1.2k GitHub stars~812 tokensUpdated today
    Backend & APIsAuto-check passed

More from NVIDIA/Megatron-LM

All 14 skills in this repo
  • Official

    Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.

    18k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Official

    Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.

    18k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Megatron-LM CI/CD Guide

    NVIDIA/Megatron-LM

    Official

    Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.

    18k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Official

    Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.

    18k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Official

    Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document.

    18k GitHub stars~1.6k tokensUpdated today
    Auto-check passed

Categories

Questions about Megatron Core Testing Guide

What does Megatron Core Testing Guide do?

Guide to the Megatron-LM test system: layout, recipe YAML, running and adding unit and functional tests, golden values, marker filters and CI parity. The guide opens with answer-first facts on disabling tests without deleting them. In functional recipe YAML you suffix the scope with -broken, for example turning mr-github into mr-github-broken, while unit tests use pytest markers: flaky_in_dev skips in the default dev environment and flaky skips in LTS.

When should I use Megatron Core Testing Guide?

Megatron Core Testing Guide fits situations like: adding a unit or functional test to Megatron-LM; temporarily disabling a flaky test without deleting it; reproducing a CI test failure locally and checking CI parity.

How do I install Megatron Core Testing Guide in Claude Code?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-testing -a claude-code`. Or copy the skill folder (skills/mcore-testing in NVIDIA/Megatron-LM) into .claude/skills/mcore-testing in your project. Claude Code loads it when a task matches its description.

How do I install Megatron Core Testing Guide in Codex?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-testing -a codex`. Or copy the skill folder (skills/mcore-testing in NVIDIA/Megatron-LM) into .agents/skills/mcore-testing in your project. Codex loads it when a task matches its description.

Can I use Megatron Core Testing Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/Megatron-LM --skill mcore-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mcore-testing, .gemini/skills/mcore-testing, .github/skills/mcore-testing and .opencode/skills/mcore-testing in your project.

What does Megatron Core Testing Guide need to run?

Going by SKILL.md and its folder, Megatron Core Testing Guide needs the command-line tools its instructions call (uv and python). Our summary lists: A Megatron-LM checkout; GPU access for unit tests; Docker for CI-style runs.

Does Megatron Core Testing Guide access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Megatron Core Testing Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Megatron Core Testing Guide use?

Megatron Core Testing Guide is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Megatron Core Testing Guide use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Megatron Core Testing Guide?

Skills that share tags, products or a category with Megatron Core Testing Guide: Simple Modern Uv (jlevy/simple-modern-uv, 301 stars), Robotics Testing (arpitg1304/robotics-agent-skills, 368 stars), Sitl Testing (ArduPilot/MethodicConfigurator, 163 stars) and Gating Deid Leakage (maziyarpanahi/openmed, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Megatron Core Testing Guide?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/Megatron-LM, which has 18,078 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/Megatron-LM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.