Agent skill

Megatron Checkpoint Layout

by EvolvingLMMs-Lab in EvolvingLMMs-Lab/LLaVA-OneVision-2

Bilingual guidance for Megatron checkpoint 1D 2D 3D mprank layouts across tensor pipeline and expert parallel dimensions

Apache-2.0Auto-check passedWriting & Content

Install Megatron Checkpoint Layout

skills CLI
$ npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill megatron-checkpoint-layout -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvolvingLMMs-Lab/LLaVA-OneVision-2 megatron-checkpoint-layout --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.opencode/skills/megatron-checkpoint-layout .claude/skills/megatron-checkpoint-layout && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
megatron-checkpoint-layout
GitHub stars
1.2k
Token cost
~1k tokens
SKILL.md length
532 words
Files
1
Skills in repo
8
Repo updated
First seen
Licence
Apache-2.0

At a glance

Bilingual guidance for Megatron checkpoint 1D 2D 3D mprank layouts across tensor pipeline and expert parallel dimensions

  • Works in 8 steps: First decide whether EP exists in the… → If EP exists, require… → If EP does not exist, read as… → …
  • Tasks that involve Translation
  • SKILL.md covers Purpose / 用途, Core rule / 核心规则, Mental model / 心智模型 and Practical interpretation /…, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Megatron Checkpoint Layout is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2. Bilingual guidance for Megatron checkpoint 1D 2D 3D mprank layouts across tensor pipeline and expert parallel dimensions

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: opencode

It sits in Writing & Content, covering Translation. It works with NVIDIA AI Platform. The repository describes itself as: Fully Open Framework for Democratized Multimodal Training. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Translation

Example prompts

  • “/megatron-checkpoint-layout”

Requirements

  • Compatibility (from SKILL.md): opencode

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. First decide whether EP exists in the checkpoint contract.
  2. If EP exists, require mp_rank_{tp}_{pp}_{ep}.
  3. If EP does not exist, read as mp_rank_{tp} or mp_rank_{tp}_{pp}.
  4. Do not infer 3D solely from pipeline_model_parallel_size > 1.
  5. 先判断这个 checkpoint 契约里是否存在 EP。
  6. 如果存在 EP,就要求目录是 mp_rank_{tp}_{pp}_{ep}。
  7. 如果不存在 EP,就按 mp_rank_{tp} 或 mp_rank_{tp}_{pp} 去读。
  8. 不要仅凭 pipeline_model_parallel_size > 1 就推断它一定是 3D。

What it can do on your machine

Read from SKILL.md and the folder at commit 6ef16b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    opencode

    From compatibility in the SKILL.md frontmatter.

Context cost

Megatron Checkpoint Layout loads about 1k tokens when it runs. Until then it costs about 37 tokens; SKILL.md has 532 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~37
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvolvingLMMs-Lab/LLaVA-OneVision-2 at commit 6ef16b1, republished under its Apache-2.0 licence (© EvolvingLMMs-Lab). 532 words, ~1,017 tokens.

Download SKILL.mdSave it as .claude/skills/megatron-checkpoint-layout/SKILL.md (or your agent's skills folder).
name
megatron-checkpoint-layout
description
Bilingual guidance for Megatron checkpoint 1D 2D 3D mp_rank layouts across tensor pipeline and expert parallel dimensions
compatibility
opencode
metadata.domain
distributed-training
metadata.framework
megatron
metadata.repo
llava-onevision2

Purpose / 用途

Use this skill when diagnosing, designing, or converting Megatron/Megatron-Core checkpoints that may use TP, PP, and EP.

在排查、设计或转换使用 TP、PP、EP 的 Megatron / Megatron-Core checkpoint 时,使用这个 skill。

Core rule / 核心规则

  • TP only: mp_rank_{tp}

  • TP + PP: mp_rank_{tp}_{pp}

  • TP + PP + EP: mp_rank_{tp}_{pp}_{ep}

  • 只有 TP:mp_rank_{tp}

  • TP + PP:mp_rank_{tp}_{pp}

  • TP + PP + EP:mp_rank_{tp}_{pp}_{ep}

The key discriminator is whether expert parallelism participates in checkpoint sharding.

真正的分界点是:expert parallelism 是否参与了 checkpoint 切分。

  • If EP is present, treat the checkpoint layout as 3D.

  • If EP is absent, treat the checkpoint layout as non-EP and use 1D or 2D.

  • 如果存在 EP,就按 3D 布局处理。

  • 如果不存在 EP,就按非 EP 布局处理,即 1D 或 2D。

Mental model / 心智模型

Megatron does not treat pp > 1 as meaning 3D by itself.

Megatron 不会因为 pp > 1 就自动把 checkpoint 视为 3D。

  • PP adds a pipeline index.

  • EP adds an expert index.

  • The third coordinate exists because EP exists, not because PP exists.

  • PP 只是在目录里增加 pipeline 这一维。

  • EP 才会增加 expert 这一维。

  • 第三维存在的原因是 EP 存在,而不是因为 PP 存在。

So even if tp=1 and pp=1, once EP is enabled the checkpoint naming is still conceptually 3D because ranks are addressed by (tp, pp, ep).

所以即使 tp=1 且 pp=1,只要启用了 EP,checkpoint 在语义上仍然是 3D,因为 rank 仍然由 (tp, pp, ep) 共同定位。

Practical interpretation / 实际使用解释

When reading or converting checkpoints:

在读取或转换 checkpoint 时:

  1. First decide whether EP exists in the checkpoint contract.

  2. If EP exists, require mp_rank_{tp}_{pp}_{ep}.

  3. If EP does not exist, read as mp_rank_{tp} or mp_rank_{tp}_{pp}.

  4. Do not infer 3D solely from pipeline_model_parallel_size > 1.

  5. 先判断这个 checkpoint 契约里是否存在 EP。

  6. 如果存在 EP,就要求目录是 mp_rank_{tp}_{pp}_{ep}。

  7. 如果不存在 EP,就按 mp_rank_{tp} 或 mp_rank_{tp}_{pp} 去读。

  8. 不要仅凭 pipeline_model_parallel_size > 1 就推断它一定是 3D。

Typical failure pattern / 典型错误模式

Bad assumption:

错误假设:

  • pp > 1 so loader chooses a 3D reader.

  • 只要 pp > 1,loader 就应该走 3D reader。

Why it fails:

为什么会失败:

  • Dense non-MoE checkpoints with TP+PP usually use mp_rank_{tp}_{pp} only.

  • A 3D loader then looks for an EP coordinate that is not present.

  • 普通 dense、非 MoE 的 TP+PP checkpoint,通常只有 mp_rank_{tp}_{pp}。

  • 这时 3D loader 会去找并不存在的 EP 坐标,最终报错。

Show full SKILL.md (210 more words)Show less

For this repository, follow this rule:

这个仓库建议遵循以下规则:

  • if expert_parallel_size is passed, require 3D

  • if expert_parallel_size is not passed, use non-EP loading

  • 如果传了 expert_parallel_size,就强制按 3D 处理

  • 如果没有传 expert_parallel_size,就走非 EP 的加载逻辑

This matches Megatron's path-building logic better than using pp > 1 as the branch condition.

这个规则比“用 pp > 1 作为分支条件”更贴近 Megatron 自己的路径生成逻辑。

What to check during debugging / 调试时要检查什么

  • What are the actual shard directory names under the checkpoint root?

  • Was expert_parallel_size provided by the caller?

  • Is the model dense or MoE?

  • Is the loader branching on EP or incorrectly branching on PP?

  • If conversion failed, which exact mp_rank_* pattern was expected and which one exists on disk?

  • checkpoint 根目录下,实际 shard 目录名是什么?

  • 调用方是否传入了 expert_parallel_size?

  • 当前模型是 dense 还是 MoE?

  • loader 是按 EP 分支,还是错误地按 PP 分支?

  • 如果转换失败,程序期望的 mp_rank_* 模式是什么,磁盘上实际又是什么?

Expected outputs when using this skill / 使用本 skill 时的期望输出

When asked to analyze a checkpoint issue, return:

当你被要求分析 checkpoint 问题时,应该返回:

  1. the inferred layout class: 1D, 2D, or 3D

  2. the reason for that classification

  3. the expected directory naming pattern

  4. whether the caller should use a non-EP loader or an EP-aware loader

  5. any mismatch between runtime arguments and on-disk shard layout

  6. 推断出的布局类别:1D、2D 或 3D

  7. 这样分类的原因

  8. 期望的目录命名模式

  9. 调用方应使用非 EP loader 还是 EP-aware loader

  10. 运行时参数与磁盘上 shard 布局之间是否存在不匹配

© EvolvingLMMs-Lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .opencode/skills/megatron-checkpoint-layout of EvolvingLMMs-Lab/LLaVA-OneVision-2.

Open the folder on GitHubat commit 6ef16b1

Compare with similar skills

Megatron Checkpoint Layout next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Megatron Checkpoint Layout compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Megatron Checkpoint Layout this skillEvolvingLMMs-Lab/LLaVA-OneVision-21.2k—~1kAutomated safety check: PassApache-2.0
Nemo Fabric IntegrateNVIDIA/skills3.5k—~5.5kAutomated safety check: PassApache-2.0
Translation Diff ExportDevolutions/UniGetUI26k—~1.1kAutomated safety check: PassMIT
Sync Translationssymfony/symfony31k—~1.9kAutomated safety check: PassMIT
Translation Diff ImportDevolutions/UniGetUI26k—~750Automated safety check: PassMIT
Translation Diff TranslateDevolutions/UniGetUI26k—~934Automated safety check: PassMIT

Similar skills

  • Official

    A skill your agent uses when integrating NVIDIA NeMo Fabric into a consumer application, service, evaluation harness, or platform through the typed Python SDK — translating the consumer's own…

    3.5k GitHub stars~5.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Translation Diff Export

    Devolutions/UniGetUI

    Compares UniGetUI JSON locale files against English, identifies untranslated or source-changed keys, and generates patch, reference, and handoff files for a target language.

    26k GitHub stars~1.1k tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Sync Translations

    symfony/symfony

    Synchronize translation catalogs across maintained Symfony branches: find messages that newer branches added to the English catalogs but that are still missing from the oldest maintained branch…

    31k GitHub stars~1.9k tokensUpdated today
    Writing & ContentAuto-check passed
  • Translation Diff Import

    Devolutions/UniGetUI

    Merges translated key-value pairs from a UniGetUI JSON localization patch back into the full language file and validates the merged result.

    26k GitHub stars~750 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Translation Diff Translate

    Devolutions/UniGetUI

    Translates a sparse UniGetUI JSON language patch, writes completed entries into the working copy, preserves placeholders and terminology, and prepares the patch for merge-back.

    26k GitHub stars~934 tokensUpdated yesterday
    Writing & ContentAuto-check passed
  • Generate Translations

    payloadcms/payload

    A skill your agent uses when new translation keys are added to packages to generate new translations strings

    45k GitHub stars~1.1k tokensUpdated today
    Writing & ContentAuto-check passed

More from EvolvingLMMs-Lab/LLaVA-OneVision-2

All 8 skills in this repo
  • Commit Message

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Guide for writing clear, consistent git commit messages following this repository's conventions

    1.2k GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed
  • Cu Lengths Attention Flow

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding how culengths controls attention behavior across ViT and LLM stages, and how patchpositions scope differs between the two

    1.2k GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Distributed Offline Packing

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running offlinepacking/autopipe.sh across multiple nodes to produce padding-free packed WebDataset shards for SFT, with Energon Metadataset assembly

    1.2k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Length Pool Sort Dataset

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for understanding LengthPoolSortDataset cross-rank length synchronization mechanism in multi-GPU training

    1.2k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Llava Onevision2 Consistency

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for running and interpreting LLaVA-OneVision2 HF vs Megatron consistency checks across TP and PP settings

    1.2k GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Offline Packing Env Vars

    EvolvingLMMs-Lab/LLaVA-OneVision-2

    Bilingual guide for the OFFLINEPACKINGBMR and OFFLINEPACKEDDATA environment variables that control LLaVA-OneVision2 training-side packing — what each gate does, why both must be enabled together…

    1.2k GitHub stars~3.9k tokensUpdated yesterday
    Auto-check passed

Questions about Megatron Checkpoint Layout

What does Megatron Checkpoint Layout do?

Bilingual guidance for Megatron checkpoint 1D 2D 3D mprank layouts across tensor pipeline and expert parallel dimensions. Megatron Checkpoint Layout is an agent skill from EvolvingLMMs-Lab/LLaVA-OneVision-2.

When should I use Megatron Checkpoint Layout?

Megatron Checkpoint Layout fits situations like: tasks that involve Translation.

How do I install Megatron Checkpoint Layout in Claude Code?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill megatron-checkpoint-layout -a claude-code`. Or copy the skill folder (.opencode/skills/megatron-checkpoint-layout in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .claude/skills/megatron-checkpoint-layout in your project. Claude Code loads it when a task matches its description.

How do I install Megatron Checkpoint Layout in Codex?

Run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill megatron-checkpoint-layout -a codex`. Or copy the skill folder (.opencode/skills/megatron-checkpoint-layout in EvolvingLMMs-Lab/LLaVA-OneVision-2) into .agents/skills/megatron-checkpoint-layout in your project. Codex loads it when a task matches its description.

Can I use Megatron Checkpoint Layout in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvolvingLMMs-Lab/LLaVA-OneVision-2 --skill megatron-checkpoint-layout -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/megatron-checkpoint-layout, .gemini/skills/megatron-checkpoint-layout, .github/skills/megatron-checkpoint-layout and .opencode/skills/megatron-checkpoint-layout in your project.

What does Megatron Checkpoint Layout need to run?

SKILL.md names no scripts, command-line tools or credentials: Megatron Checkpoint Layout is instructions for the agent only. Compatibility (from SKILL.md): opencode.

Does Megatron Checkpoint Layout access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Megatron Checkpoint Layout safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Megatron Checkpoint Layout use?

Megatron Checkpoint Layout is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Megatron Checkpoint Layout use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Megatron Checkpoint Layout?

Skills that share tags, products or a category with Megatron Checkpoint Layout: Nemo Fabric Integrate (NVIDIA/skills, 3.5k stars), Translation Diff Export (Devolutions/UniGetUI, 26k stars), Sync Translations (symfony/symfony, 31k stars) and Translation Diff Import (Devolutions/UniGetUI, 26k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Megatron Checkpoint Layout?

EvolvingLMMs-Lab (a GitHub organization) maintains it in EvolvingLMMs-Lab/LLaVA-OneVision-2, which has 1,216 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: EvolvingLMMs-Lab/LLaVA-OneVision-2 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.