Official agent skill

Nemo Mbridge Perf Hierarchical Context Parallel

by NVIDIA in NVIDIA/skills

Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

OfficialApache-2.0Auto-check passed

Install Nemo Mbridge Perf Hierarchical Context Parallel

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-hierarchical-context-parallel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-hierarchical-context-parallel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-hierarchical-context-parallel .claude/skills/nemo-mbridge-perf-hierarchical-context-parallel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-hierarchical-context-parallel
GitHub stars
3.5k
Token cost
~1.4k tokens
SKILL.md length
328 words
Files
6
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

  • Works in 6 steps: Bridge HCP is MPU-only today: If… → No checked-in Bridge recipe currently… → Single-GPU load helpers clear… → …
  • SKILL.md covers Enablement, Code Anchors, Implementation Map and Pitfalls, plus 1 more section
  • Calls uv

What it does

Nemo Mbridge Perf Hierarchical Context Parallel is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

Example prompts

  • “/nemo-mbridge-perf-hierarchical-context-parallel”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Bridge HCP is MPU-only today: If use_decentralized_pg=True, Bridge initializes flat CP groups and leaves HCP unset.
  2. No checked-in Bridge recipe currently exercises HCP directly.
  3. Single-GPU load helpers clear hierarchical_context_parallel_sizes.
  4. Silent broken training on old stacks: If you use a2a+p2p without setting hierarchical_context_parallel_sizes, MCore now asserts. Older…
  5. Product must match: prod(hierarchical_context_parallel_sizes) must exactly equal context_parallel_size. A mismatch triggers an assertion.
  6. Verify in logs: Look for the process group initialization output. You should see HIERARCHICAL_CONTEXT_PARALLEL_GROUPS being created. If…

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Hierarchical Context Parallel loads about 1.4k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 328 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 328 words, ~1,411 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-hierarchical-context-parallel/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-hierarchical-context-parallel
description
Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.
license
Apache-2.0
when_to_use
Scaling context parallelism beyond KV heads, or investigating a commit that changed CP config and caused OOM or a regression…

Hierarchical Context Parallel Skill

This skill covers hierarchical context parallelism: nested context-parallel process groups used by cp_comm_type="a2a+p2p" and configured with hierarchical_context_parallel_sizes.

For what hierarchical CP is, when to use it, and the decision tree (a2a+p2p vs pure a2a vs p2p), see:

  • @docs/training/hierarchical-context-parallel.md
  • @skills/nemo-mbridge-perf-hierarchical-context-parallel/card.yaml

Enablement

Minimal Bridge override:

python
cfg.model.context_parallel_size = 4
cfg.model.cp_comm_type = "a2a+p2p"
cfg.model.hierarchical_context_parallel_sizes = [2, 2]
cfg.dist.use_decentralized_pg = False

Required constraints:

  • prod(hierarchical_context_parallel_sizes) == context_parallel_size
  • seq_length % (2 * context_parallel_size) == 0
  • Transformer Engine >= 1.12.0

Code Anchors

Upstream config and validation:

45543rdpartyMegatron
context_parallel_size: int = 1
"""Splits network input along sequence dimension across GPU ranks."""

hierarchical_context_parallel_sizes: Optional[list[int]] = None
"""Degrees of the hierarchical context parallelism. Users should provide a list to specify 
   the sizes for different levels. Taking the a2a+p2p cp comm type as example, it contains
   groups of two levels, so the first value of the list indicates the group size of the a2a
   communication type, and the second value indicates the group size of the p2p communication
   type.
"""
4284333rdpartyMegatr
if args.hierarchical_context_parallel_sizes:
    from numpy import prod
    assert args.context_parallel_size == prod(args.hierarchical_context_parallel_sizes)
if "a2a+p2p" in args.cp_comm_type:
    assert args.hierarchical_context_parallel_sizes is not None, \
    "--hierarchical-context-parallel-sizes must be set when a2a+p2p is used in cp comm"

Bridge MPU path:

613648srcmegatronbri
parallel_state.initialize_model_parallel(
    ...
    context_parallel_size=model_config.context_parallel_size,
    hierarchical_context_parallel_sizes=model_config.hierarchical_context_parallel_sizes,
    ...
)
...
return ProcessGroupCollection.use_mpu_process_groups()

Bridge decentralized-PG path:

503524srcmegatronbri
pg_collection = ProcessGroupCollection(
    ...
    cp=cp_pg,
    tp_cp=tp_cp_pg,
    hcp=None,
    ep=ep_pg,
    ...
)

Implementation Map

The code anchors above show the config declarations and argument validation.

Validation (MCore)

TransformerConfig.__post_init__ enforces that a2a+p2p requires HCP sizes and the product matches CP.

Process group creation

parallel_state.initialize_model_parallel creates hierarchical CP sub-groups when HCP sizes are provided via create_hierarchical_groups. Bridge currently gets those groups through the MPU-backed ProcessGroupCollection.

TE integration

TEDotProductAttention passes the hierarchical groups to Transformer Engine when a2a+p2p is used. Requires Transformer Engine >= 1.12.0.

Pitfalls

  1. Bridge HCP is MPU-only today: If use_decentralized_pg=True, Bridge initializes flat CP groups and leaves HCP unset.
  2. No checked-in Bridge recipe currently exercises HCP directly.
  3. Single-GPU load helpers clear hierarchical_context_parallel_sizes.
  4. Silent broken training on old stacks: If you use a2a+p2p without setting hierarchical_context_parallel_sizes, MCore now asserts. Older versions would silently disable CP communication, so each rank attended only to its local chunk and produced artificially high throughput with broken gradients.
  5. Product must match: prod(hierarchical_context_parallel_sizes) must exactly equal context_parallel_size. A mismatch triggers an assertion.
  6. Verify in logs: Look for the process group initialization output. You should see HIERARCHICAL_CONTEXT_PARALLEL_GROUPS being created. If you only see CONTEXT_PARALLEL_GROUP, HCP is not active.

Verification

No dedicated Bridge end-to-end test exists yet for HCP (see @skills/nemo-mbridge-perf-hierarchical-context-parallel/card.yaml follow_up_validation). Use the existing unit tests and log inspection instead.

Run the decentralized-PG unit test to confirm the flat-CP behavior is preserved:

bash
uv run python -m pytest tests/unit_tests/training/test_decentralized_pg.py -q

For a manual smoke check, launch a 4-GPU run with a small recipe and cp_comm_type=a2a+p2p plus hierarchical_context_parallel_sizes=[2,2]:

bash
CUDA_VISIBLE_DEVICES=0,1,2,3 uv run python -m torch.distributed.run --nproc_per_node=4 \
  scripts/training/run_recipe.py \
  --recipe llama32_1b_pretrain_config \
  model.context_parallel_size=4 \
  model.cp_comm_type=a2a+p2p \
  "model.hierarchical_context_parallel_sizes=[2,2]" \
  train.train_iters=2

Success criteria:

  • Logs show HIERARCHICAL_CONTEXT_PARALLEL_GROUPS being created
  • Training completes at least one step without error
  • If you only see CONTEXT_PARALLEL_GROUP, HCP is not active

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-hierarchical-context-parallel of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemo Mbridge Perf Hierarchical Context Parallel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Hierarchical Context Parallel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Hierarchical Context Parallel this skillNVIDIA/skills3.5k—~1.4kAutomated safety check: PassApache-2.0
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Skill InspectorNVIDIA/SkillSpector20k—~1.8kAutomated safety check: PassApache-2.0
Megatron-LM Container and Dependency SetupNVIDIA/Megatron-LM18k—~2.6kAutomated safety check: PassApache-2.0
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
Megatron-LM Base Image BumpNVIDIA/Megatron-LM18k—~2.8kAutomated safety check: PassApache-2.0

Similar skills

  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Skill Inspector

    NVIDIA/SkillSpector

    Official

    Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.

    20k GitHub stars~1.8k tokensUpdated today
    SecurityAuto-check passed
  • Official

    Walks an agent through working inside the Megatron-LM CI container and changing dependencies with uv, so lock files resolve the same locally and in CI.

    18k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM Base Image Bump

    NVIDIA/Megatron-LM

    Official

    Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up.

    18k GitHub stars~2.8k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Remove bracketed NemoClaw tags from GitHub issue and PR titles.

    23k GitHub stars~693 tokensUpdated today
    Marketing & SEOAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Nemo Mbridge Perf Hierarchical Context Parallel

What does Nemo Mbridge Perf Hierarchical Context Parallel do?

Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification. Nemo Mbridge Perf Hierarchical Context Parallel is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Operational guide for enabling hierarchical context parallelism in Megatron-Bridge, including config knobs, code anchors, pitfalls, and verification.

How do I install Nemo Mbridge Perf Hierarchical Context Parallel in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-hierarchical-context-parallel -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-hierarchical-context-parallel in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-hierarchical-context-parallel in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Hierarchical Context Parallel in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-hierarchical-context-parallel -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-hierarchical-context-parallel in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-hierarchical-context-parallel in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Hierarchical Context Parallel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-hierarchical-context-parallel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-hierarchical-context-parallel, .gemini/skills/nemo-mbridge-perf-hierarchical-context-parallel, .github/skills/nemo-mbridge-perf-hierarchical-context-parallel and .opencode/skills/nemo-mbridge-perf-hierarchical-context-parallel in your project.

What does Nemo Mbridge Perf Hierarchical Context Parallel need to run?

Going by SKILL.md and its folder, Nemo Mbridge Perf Hierarchical Context Parallel needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Perf Hierarchical Context Parallel access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Hierarchical Context Parallel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Hierarchical Context Parallel use?

Nemo Mbridge Perf Hierarchical Context Parallel is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Hierarchical Context Parallel use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Hierarchical Context Parallel?

Skills that share tags, products or a category with Nemo Mbridge Perf Hierarchical Context Parallel: LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Skill Inspector (NVIDIA/SkillSpector, 20k stars), Megatron-LM Container and Dependency Setup (NVIDIA/Megatron-LM, 18k stars) and Embeddings via 9Router (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Hierarchical Context Parallel?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.