Official agent skill

Megatron GPT to Hybrid Migration

by NVIDIA in NVIDIA/Megatron-LM

Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Megatron GPT to Hybrid Migration

skills CLI
$ npx skills add NVIDIA/Megatron-LM --skill mcore-migrate-gpt-to-hybrid -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/Megatron-LM mcore-migrate-gpt-to-hybrid --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/Megatron-LM.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mcore-migrate-gpt-to-hybrid .claude/skills/mcore-migrate-gpt-to-hybrid && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mcore-migrate-gpt-to-hybrid
GitHub stars
18k
Token cost
~1.6k tokens
SKILL.md length
583 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document.

  • Works in 5 steps: Pull the task artifact first: checkpoint… → Read the canonical migration document… → Follow only the relevant document… → …
  • Converting a GPTModel checkpoint or provider to HybridModel
  • SKILL.md covers Answer-First Migration Guidance, Workflow, Transferring an Existing… and Documentation Drift
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

This skill covers migrating a Megatron Core GPTModel setup to HybridModel: checkpoints, model providers, training commands and layer mappings. It treats docs/user-guide/hybrid-model-migration.md in the repository as the canonical source, tells the agent to read it completely before answering, planning, editing, converting or training, and keeps migration behavior in that document rather than duplicating it. The skill adds only the mechanical procedure for editing an existing pretrain_gpt.py launch script and the shell hazards that come with it.

The workflow starts by collecting the task artifact (checkpoint metadata, config, training command, conversion log, diff or failure output), follows only the relevant document sections without inventing an unsupported path or silently changing the target architecture, validates in proportion, and reports with a link to the document. Launch script edits swap pretrain_gpt.py for pretrain_hybrid.py, generate the layer pattern with a shell snippet instead of typing it by hand, since a 96-layer model needs a 192-character pattern, replace --num-layers with --hybrid-layer-pattern, because a stale value only warns while being silently overridden, and add the stack spec in place of any GPT spec.

When your agent uses it

  • Converting a GPTModel checkpoint or provider to HybridModel
  • Editing a pretrain_gpt.py launch script into its hybrid equivalent
  • Mapping GPT layers to a hybrid layer pattern for dense or MoE blocks
  • Debugging a migrated training command that behaves unexpectedly

Example prompts

  • “Migrate my pretrain_gpt.py launch script for a 32-layer GPT model to the hybrid entrypoint.”
  • “Build the hybrid layer pattern for a 32-layer dense model split into 4 pipeline segments.”
  • “My converted run starts fine, but I suspect a stale --num-layers value is being overridden, so check the launch command.”

Requirements

  • A Megatron-LM checkout that includes docs/user-guide/hybrid-model-migration.md

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Pull the task artifact first: checkpoint metadata, model provider or config,
  2. Read the canonical migration document completely.
  3. Follow only the relevant document sections. Do not invent an unsupported
  4. Validate the result proportionately, invoking the relevant repository build
  5. Report the outcome and link the canonical document for human readers.

What it can do on your machine

Read from SKILL.md and the folder at commit a3e1f82. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Megatron GPT to Hybrid Migration loads about 1.6k tokens when it runs. Until then it costs about 173 tokens; SKILL.md has 583 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~173
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/Megatron-LM at commit a3e1f82, republished under its Apache-2.0 licence (© NVIDIA). 583 words, ~1,636 tokens.

Download SKILL.mdSave it as .claude/skills/mcore-migrate-gpt-to-hybrid/SKILL.md (or your agent's skills folder).
name
mcore-migrate-gpt-to-hybrid
description
Migration guide for moving Megatron Core GPTModel checkpoints, model providers, training commands, and layer mappings to HybridModel, including the mechanical steps for transferring an existing pretrain_gpt.py launch script. Use when migrating or reviewing a GPTModel checkpoint or training workflow for HybridModel; transferring an existing pretrain_gpt.py script, sbatch, or launcher to pretrain_hybrid.py; choosing or reviewing a hybrid layer pattern; running gpt_hybrid_conversion.py; loading a converted checkpoint; diagnosing GPT-to-Hybrid migration issues; 'migrate GPTModel to HybridModel', 'convert GPT checkpoint to HybridModel', 'hybrid layer pattern'.
license
Apache-2.0
metadata.author
Philip Petrakian <ppetrakian@nvidia.com>

GPTModel to HybridModel Migration

Answer-First Migration Guidance

  • The canonical source is docs/user-guide/hybrid-model-migration.md.
  • Read the canonical document completely before answering, planning, reviewing, editing, converting, or training.
  • Keep migration behavior, commands, mappings, prerequisites, limitations, and validation in the canonical document only. Do not duplicate them in this skill.
  • This skill adds only what the canonical document does not cover: the mechanical procedure for editing an existing launch script, and the shell hazards that procedure runs into.

Workflow

  1. Pull the task artifact first: checkpoint metadata, model provider or config, training command, conversion log, diff, or failure output.
  2. Read the canonical migration document completely.
  3. Follow only the relevant document sections. Do not invent an unsupported migration path or silently change the target architecture.
  4. Validate the result proportionately, invoking the relevant repository build and testing skills when applicable.
  5. Report the outcome and link the canonical document for human readers.

Transferring an Existing Launch Script

The canonical document specifies what the migrated command must contain. This section covers how to edit a working script into it without silent breakage. Apply the edits in order.

1. Entrypoint. pretrain_gpt.py → pretrain_hybrid.py. When the entrypoint comes from a shell variable or a wrapper, follow it to the real invocation.

2. Generate the pattern instead of typing it. A 96-layer model needs a 192-character pattern; hand-typing invites a silent off-by-one.

bash
n=32                                   # source GPT layer count
blk='*-'                               # '*-' dense, '*E' every-layer MoE
pat=$(printf "$blk%.0s" $(seq $n))

With pipeline segments (seg must divide n, and the segment count must be divisible by --pipeline-model-parallel-size):

bash
n=32; seg=4; per=$((n/seg))
b=$(printf "$blk%.0s" $(seq $per)); pat=$b
for ((i=1;i<seg;i++)); do pat="$pat|$b"; done

3. Replace --num-layers N with --hybrid-layer-pattern. Deleting --num-layers matters: leaving a stale value is only a warning, so it looks healthy while being silently overridden by the pattern-derived count.

4. Add the stack spec, replacing any GPT --spec rather than adding a second one.

5. Delete the pipeline-layout arguments the parser rejects, and repoint --save at a new directory. See the canonical document for both lists.

Show full SKILL.md (277 more words)Show less
Bash-array scripts: quoting at the definition site is not enough

Most scripts under examples/ collect arguments in arrays and expand them unquoted:

bash
torchrun ${DISTRIBUTED_ARGS[@]} pretrain_gpt.py ${MODEL_ARGS[@]}

Unquoted ${ARR[@]} re-runs word-splitting and pathname expansion on every element, so the pattern is globbed against the launch directory at expansion time — single-quoting it where the array is defined does not protect it:

bash
touch 'a-b-'; ARGS=(--hybrid-layer-pattern '*-*-')
printf '[%s]\n' ${ARGS[@]}     # -> [--hybrid-layer-pattern] [a-b-]   silently corrupted
printf '[%s]\n' "${ARGS[@]}"   # -> [--hybrid-layer-pattern] [*-*-]   correct

An unmatched glob survives intact, so this passes by luck in most working directories and fails only when some file happens to match. Store the pattern in a variable and quote that array's expansion:

bash
HYBRID_PATTERN=$(printf '*-%.0s' $(seq $NUM_LAYERS))
MODEL_ARGS=( ... --hybrid-layer-pattern "$HYBRID_PATTERN" ... )
torchrun "${DISTRIBUTED_ARGS[@]}" pretrain_hybrid.py "${MODEL_ARGS[@]}" ...
Verify the rewrite

Both checks are cheap and catch the common slips:

bash
# 1. No rejected or stale arguments survived -- must print nothing.
grep -nE -- '--(num-layers|num-layers-per-virtual-pipeline-stage|num-virtual-stages-per-pipeline-rank|pipeline-model-parallel-layout|account-for-embedding-in-pipeline-split|account-for-loss-in-pipeline-split|hybrid-override-pattern|fim-data)\b' train_hybrid.sh

# 2. Pattern shape -- attn and mlp must each equal the source GPT layer count.
p='*-*-|*-*-'
main=${p%%/*}; main=${main//|/}
attn=${main//[^\*]/}; mlp=${main//[^-E]/}; segs=${p%%/*}; segs=${segs//[^|]/}
echo "layers=${#main} attn=${#attn} mlp=${#mlp} segments=$(( ${#segs} + 1 ))"

Then diff the migrated script against the original: it should contain the edits above and nothing else.

Expected result of an architecture-preserving transfer

A *- or *E transfer changes the layer indexing, not the model. On a measured 8-block dense run (2 GPUs, bf16, seq 4096, 100 iterations, identical seed and data), pretrain_gpt.py --num-layers 8 and pretrain_hybrid.py --hybrid-layer-pattern '*-*-*-*-*-*-*-*-' produced:

  • identical parameter counts (2,818,641,920 on both);
  • HybridModel: ... layers='*-*-*-*-*-*-*-*-' (16 layers) from the allocator;
  • steady-state throughput within 0.1% (490.7 vs 491.2 ms/iter);
  • identical loss for the first two iterations, then a zero-mean drift of |Δ| ≤ 0.08 attributable to kernel/reduction ordering.

Treat a systematic loss offset, a parameter-count difference, or a throughput gap beyond noise as a migration bug, not as expected behavior. Note that per-iteration wall clock early in a run is dominated by dataset-cache warmup, so compare steady-state iterations only.


Documentation Drift

If the implementation and migration guide disagree:

  1. Report the discrepancy before continuing.
  2. If the task authorizes a correction, update the canonical document first.
  3. Do not add a competing migration rule to this skill.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/mcore-migrate-gpt-to-hybrid of NVIDIA/Megatron-LM.

Open the folder on GitHubat commit a3e1f82

Compare with similar skills

Megatron GPT to Hybrid Migration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Megatron GPT to Hybrid Migration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Megatron GPT to Hybrid Migration this skillNVIDIA/Megatron-LM18k—~1.6kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
Megatron-Core LLM TrainingOrchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT
Quark Env Preflightamd/Quark182—~1.4kAutomated safety check: PassMIT
Slime Useryzlnew/infra-skills149—~3.2kAutomated safety check: PassNone
Hyperpod Version Checkerawslabs/agent-plugins916—~910Automated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Megatron-Core LLM Training

    Orchestra-Research/AI-Research-SKILLs

    Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count.

    13k GitHub starsUsed in 2 repos~2.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    182 GitHub stars~1.4k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check passed
  • Slime User

    yzlnew/infra-skills

    Guide for using SLIME (LLM post-training framework for RL Scaling).

    149 GitHub stars~3.2k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    916 GitHub stars~910 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/Megatron-LM

All 14 skills in this repo
  • Official

    Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.

    18k GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Megatron-LM CI/CD Guide

    NVIDIA/Megatron-LM

    Official

    Explains Megatron-LM's CI pipeline, PR scope labels, triggering the internal GitLab CI with a dry run first, and investigating CI failures.

    18k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Official

    Investigates a failing GitHub Actions run or job for Megatron-LM, finds the root cause plus the PR and test author involved, and files a structured bug issue.

    18k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Official

    Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.

    18k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Questions about Megatron GPT to Hybrid Migration

What does Megatron GPT to Hybrid Migration do?

Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document. This skill covers migrating a Megatron Core GPTModel setup to HybridModel: checkpoints, model providers, training commands and layer mappings.md in the repository as the canonical source, tells the agent to read it completely before answering, planning, editing, converting or training, and keeps migration behavior in that document rather than duplicating it.

When should I use Megatron GPT to Hybrid Migration?

Megatron GPT to Hybrid Migration fits situations like: converting a GPTModel checkpoint or provider to HybridModel; editing a pretrain_gpt.py launch script into its hybrid equivalent; mapping GPT layers to a hybrid layer pattern for dense or MoE blocks; debugging a migrated training command that behaves unexpectedly.

How do I install Megatron GPT to Hybrid Migration in Claude Code?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-migrate-gpt-to-hybrid -a claude-code`. Or copy the skill folder (skills/mcore-migrate-gpt-to-hybrid in NVIDIA/Megatron-LM) into .claude/skills/mcore-migrate-gpt-to-hybrid in your project. Claude Code loads it when a task matches its description.

How do I install Megatron GPT to Hybrid Migration in Codex?

Run `npx skills add NVIDIA/Megatron-LM --skill mcore-migrate-gpt-to-hybrid -a codex`. Or copy the skill folder (skills/mcore-migrate-gpt-to-hybrid in NVIDIA/Megatron-LM) into .agents/skills/mcore-migrate-gpt-to-hybrid in your project. Codex loads it when a task matches its description.

Can I use Megatron GPT to Hybrid Migration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/Megatron-LM --skill mcore-migrate-gpt-to-hybrid -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mcore-migrate-gpt-to-hybrid, .gemini/skills/mcore-migrate-gpt-to-hybrid, .github/skills/mcore-migrate-gpt-to-hybrid and .opencode/skills/mcore-migrate-gpt-to-hybrid in your project.

What does Megatron GPT to Hybrid Migration need to run?

SKILL.md names no scripts, command-line tools or credentials: Megatron GPT to Hybrid Migration is instructions for the agent only. Our summary lists: A Megatron-LM checkout that includes docs/user-guide/hybrid-model-migration.md.

Does Megatron GPT to Hybrid Migration access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Megatron GPT to Hybrid Migration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Megatron GPT to Hybrid Migration use?

Megatron GPT to Hybrid Migration is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Megatron GPT to Hybrid Migration use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Megatron GPT to Hybrid Migration?

Skills that share tags, products or a category with Megatron GPT to Hybrid Migration: Graphsignal (graphsignal/graphsignal, 257 stars), Megatron-Core LLM Training (Orchestra-Research/AI-Research-SKILLs, 13k stars), Quark Env Preflight (amd/Quark, 182 stars) and Slime User (yzlnew/infra-skills, 149 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Megatron GPT to Hybrid Migration?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/Megatron-LM, which has 18,112 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 11, 2026.

Source: NVIDIA/Megatron-LM on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.