Explore Feature E2E Test
comet-ml/opik
Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.
Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-onboard-gb200-1node-tests --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mcore-onboard-gb200-1node-tests .claude/skills/mcore-onboard-gb200-1node-tests && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "mcore-onboard-gb200-1node-tests" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-onboard-gb200-1node-tests into .claude/skills/mcore-onboard-gb200-1node-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-onboard-gb200-1node-tests", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-onboard-gb200-1node-testsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-onboard-gb200-1node-tests --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/mcore-onboard-gb200-1node-tests .agents/skills/mcore-onboard-gb200-1node-tests && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "mcore-onboard-gb200-1node-tests" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-onboard-gb200-1node-tests into .agents/skills/mcore-onboard-gb200-1node-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-onboard-gb200-1node-tests", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-onboard-gb200-1node-tests --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/mcore-onboard-gb200-1node-tests .cursor/skills/mcore-onboard-gb200-1node-tests && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "mcore-onboard-gb200-1node-tests" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-onboard-gb200-1node-tests into .cursor/skills/mcore-onboard-gb200-1node-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-onboard-gb200-1node-tests", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/Megatron-LM.git --path skills/mcore-onboard-gb200-1node-tests--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-onboard-gb200-1node-tests --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/mcore-onboard-gb200-1node-tests .gemini/skills/mcore-onboard-gb200-1node-tests && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "mcore-onboard-gb200-1node-tests" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-onboard-gb200-1node-tests into .gemini/skills/mcore-onboard-gb200-1node-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-onboard-gb200-1node-tests", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/Megatron-LM mcore-onboard-gb200-1node-testsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/mcore-onboard-gb200-1node-tests .github/skills/mcore-onboard-gb200-1node-tests && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "mcore-onboard-gb200-1node-tests" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-onboard-gb200-1node-tests into .github/skills/mcore-onboard-gb200-1node-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-onboard-gb200-1node-tests", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/Megatron-LM mcore-onboard-gb200-1node-tests --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/mcore-onboard-gb200-1node-tests .opencode/skills/mcore-onboard-gb200-1node-tests && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "mcore-onboard-gb200-1node-tests" agent skill from https://github.com/NVIDIA/Megatron-LM/tree/main/skills/mcore-onboard-gb200-1node-tests into .opencode/skills/mcore-onboard-gb200-1node-tests/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mcore-onboard-gb200-1node-tests", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
mcore-onboard-gb200-1node-testsCreates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.
Each GB200 node has 4 GPUs, so a two-node test uses 8 and its one-node variant uses 4. The skill scans the products block of gpt.yaml and moe.yaml for tests scoped mr or mr-slim, skips ones already covered in the 1node recipe files and anything nightly, weekly or mr-broken, then reads each model_config.yaml for the tensor, pipeline and expert parallel sizes.
Each candidate is classified using world size as TP times PP times DP. If TP times PP is at most 4, the config is copied unchanged and DP halves on its own; if it equals 8, pipeline parallelism is halved; expert parallelism above 4 is reduced to 4; and ETP tests are re-checked afterwards. Global batch size is left alone so gradient accumulation absorbs the lower DP. New configs go in test case folders with a _1node suffix, and recipes live under tests/test_utils/recipes/gb200. The folder also carries a benchmark file, a skill card and evals.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit d5fbb65. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Megatron GB200 One-Node Test Onboarding loads about 1.3k tokens when it runs. Until then it costs about 30 tokens; SKILL.md has 486 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/Megatron-LM at commit d5fbb65, republished under its Apache-2.0 licence (© NVIDIA). 486 words, ~1,339 tokens.
.claude/skills/mcore-onboard-gb200-1node-tests/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Create 1-node (mr-github) variants of existing 2-node (mr-scoped) GB200 functional tests.
Each GB200 node has 4 GPUs. A 2-node test uses 8 GPUs total; the 1-node variant uses 4.
GB200 functional tests live in tests/test_utils/recipes/gb200/:
| Recipe file | Notes |
|---|---|
gpt.yaml | GPT dense tests, nodes: 2, gpus: 4 (8 total) |
moe.yaml | MoE tests, nodes: 2, gpus: 4 (8 total) |
moe-1node.yaml | Existing 1-node MoE tests, nodes: 1, gpus: 4 (4 total) |
gpt-1node.yaml | 1-node GPT tests (create if not present) |
Model configs live at:
tests/functional_tests/test_cases/{model}/{test_case}/model_config.yaml
1-node test cases use the _1node suffix:
tests/functional_tests/test_cases/{model}/{test_case}_1node/model_config.yaml
Scan the products: block in gpt.yaml and moe.yaml for entries with scope: [mr, ...] or scope: [mr-slim, ...]. These are the 2-node tests that need 1-node mr-github counterparts.
Ignore tests already covered in *-1node.yaml files, and ignore nightly, weekly, mr-broken scopes.
For each candidate, read its model_config.yaml and extract the key parallelism arguments:
--tensor-model-parallel-size (TP)
--pipeline-model-parallel-size (PP)
--expert-model-parallel-size (EP)
--expert-tensor-parallel-size (ETP)
--context-parallel-size (CP)
--global-batch-size
--micro-batch-sizeThe world size formula is: world_size = TP × PP × DP where DP ≥ EP.
Going from 8 GPUs → 4 GPUs:
| Condition | Action |
|---|---|
TP × PP ≤ 4 | Trivial copy. Config unchanged; DP is halved automatically. |
TP × PP = 8 (e.g. tp4 pp2) | Reduce PP. Set PP = PP / 2 (e.g. pp2→1). Verify TP × PP_new ≤ 4. |
EP > 4 (e.g. ep8 with tp1 pp1) | Reduce EP. Set EP = 4. Experts stay at num-experts (each EP rank holds more experts). |
EP > 4 and TP × PP > 4 | Reduce both PP and EP as above. |
| ETP test (ep × etp ≤ TP × DP) | Check EP × ETP ≤ TP × DP_new after PP reduction. Usually satisfied when pp→1. |
Do not change GBS — let gradient accumulation absorb the reduced DP.
_1node model config directories# Trivial copy
mkdir -p tests/functional_tests/test_cases/{model}/{test_case}_1node
cp tests/functional_tests/test_cases/{model}/{test_case}/model_config.yaml \
tests/functional_tests/test_cases/{model}/{test_case}_1node/model_config.yaml
# Then apply any parallelism changes (EP or PP) with Edit toolFor GPT tests — create tests/test_utils/recipes/gb200/gpt-1node.yaml (if absent) by cloning gpt.yaml's spec block with nodes: 1. Use this template for the spec:
type: basic
format_version: 1
maintainers: [mcore]
loggers: [stdout]
spec:
name: "{test_case}_{environment}_{platforms}"
model: gpt # or moe
build: mcore-pyt-{environment}
nodes: 1
gpus: 4
n_repeat: 5
platforms: dgx_gb200
script_setup: | # copy verbatim from gpt.yaml / moe.yaml
...
script: |- # copy verbatim from gpt.yaml / moe.yaml
...For MoE tests — append entries to the existing moe-1node.yaml.
Scope convention:
scope: [mr-github, mr-github-slim]scope: [mr-github]products:
- test_case: [<test_case>_1node]
products:
- environment: [dev]
scope: [mr-github, mr-github-slim] # or [mr-github]
platforms: [dgx_gb200]| Original (8 GPUs) | 1-node config (4 GPUs) | Notes |
|---|---|---|
| tp1 pp1 ep1 → dp8 | tp1 pp1 ep1 → dp4 | trivial |
| tp2 pp1 ep1 → dp4 | tp2 pp1 ep1 → dp2 | trivial |
| tp1 pp2 ep1 → dp4 | tp1 pp2 ep1 → dp2 | trivial |
| tp4 pp1 ep1 → dp2 | tp4 pp1 ep1 → dp1 | trivial |
| tp1 pp4 ep1 → dp2 | tp1 pp4 ep1 → dp1 | trivial |
| tp1 pp1 ep8 → dp8 | tp1 pp1 ep4 → dp4 | ep 8→4 |
| tp4 pp2 ep2 etp2 → dp1 | tp4 pp1 ep2 etp2 → dp1 | pp 2→1 |
mr-scoped tests in gpt.yaml and moe.yaml not yet in *-1node.yaml_1node/model_config.yaml for each testnodes: 1, gpus: 4mr-github scope (+ mr-github-slim for 1–2 representative tests per recipe)mr-github-slim overload (slim suite should stay small)© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files in skills/mcore-onboard-gb200-1node-tests of NVIDIA/Megatron-LM.
Open the folder on GitHubat commit d5fbb65
Megatron GB200 One-Node Test Onboarding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Megatron GB200 One-Node Test Onboarding this skillNVIDIA/Megatron-LM | 18k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Explore Feature E2E Testcomet-ml/opik | 22k | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| Kane CLI Browser TestingLambdaTest/kane-cli | 247 | — | ~8.4k | Automated safety check: Pass | Apache-2.0 | |
| Nemoclaw Maintainer Analyze CI PerformanceNVIDIA/NemoClaw | 23k | — | ~644 | Automated safety check: Pass | Apache-2.0 | |
| Endgamemicrosoft/copilot-for-eclipse | 126 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Implementnvuillam/github-dependents-info | 162 | — | ~999 | Automated safety check: Pass | MIT |
comet-ml/opik
Turns a code change into one committed, passing Playwright end-to-end spec by resolving the change scope and handing authoring to a companion skill.
LambdaTest/kane-cli
Drives a real browser through the kane-cli tool and designs requirement-linked test suites from a PRD or a plain description, with mobile and cloud-grid runs.
NVIDIA/NemoClaw
Analyze retained NemoClaw CI timings for slow CLI tests, runner queues, or base-image publication.
microsoft/copilot-for-eclipse
Orchestrate endgame verification for a GitHub milestone issue.
nvuillam/github-dependents-info
Phase 3 of the SDLC pipeline (also usable standalone). An agent skill from nvuillam/github-dependents-info.
dotnet/maui
Writes UI tests that reproduce a GitHub issue in .NET MAUI and keeps iterating until the tests actually fail, proving they catch the bug.
NVIDIA/Megatron-LM
Guide to the Megatron-LM test system: layout, recipe YAML, running and adding unit and functional tests, golden values, marker filters and CI parity.
NVIDIA/Megatron-LM
Refreshes stored golden values from a GitHub Actions run, reports signed percentage changes per model, and writes a summary ready for a pull request description.
NVIDIA/Megatron-LM
Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.
NVIDIA/Megatron-LM
Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.
NVIDIA/Megatron-LM
Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.
NVIDIA/Megatron-LM
Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.
Works with
Categories
Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs. Each GB200 node has 4 GPUs, so a two-node test uses 8 and its one-node variant uses 4.yaml for the tensor, pipeline and expert parallel sizes.
Megatron GB200 One-Node Test Onboarding fits situations like: adding one-node GitHub MR coverage for GB200 functional tests; deciding whether a two-node test config can be copied or needs smaller parallelism; creating a gpt-1node.yaml recipe when none exists.
Run `npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a claude-code`. Or copy the skill folder (skills/mcore-onboard-gb200-1node-tests in NVIDIA/Megatron-LM) into .claude/skills/mcore-onboard-gb200-1node-tests in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a codex`. Or copy the skill folder (skills/mcore-onboard-gb200-1node-tests in NVIDIA/Megatron-LM) into .agents/skills/mcore-onboard-gb200-1node-tests in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/Megatron-LM --skill mcore-onboard-gb200-1node-tests -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mcore-onboard-gb200-1node-tests, .gemini/skills/mcore-onboard-gb200-1node-tests, .github/skills/mcore-onboard-gb200-1node-tests and .opencode/skills/mcore-onboard-gb200-1node-tests in your project.
SKILL.md names no scripts, command-line tools or credentials: Megatron GB200 One-Node Test Onboarding is instructions for the agent only. Our summary lists: A checkout of the Megatron-LM repository with the GB200 test recipes.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Megatron GB200 One-Node Test Onboarding is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Megatron GB200 One-Node Test Onboarding: Explore Feature E2E Test (comet-ml/opik, 22k stars), Kane CLI Browser Testing (LambdaTest/kane-cli, 247 stars), Nemoclaw Maintainer Analyze CI Performance (NVIDIA/NemoClaw, 23k stars) and Endgame (microsoft/copilot-for-eclipse, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/Megatron-LM, which has 18,083 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.
Source: NVIDIA/Megatron-LM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.