Official agent skill

Nemo Mbridge Perf Sequence Packing

by NVIDIA in NVIDIA/skills

Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.

OfficialApache-2.0Auto-check passed

Install Nemo Mbridge Perf Sequence Packing

skills CLI
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemo-mbridge-perf-sequence-packing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-sequence-packing .claude/skills/nemo-mbridge-perf-sequence-packing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemo-mbridge-perf-sequence-packing
GitHub stars
3.5k
Token cost
~3.3k tokens
SKILL.md length
911 words
Files
6
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.

  • Works in 12 steps: Offline packed SFT, runtime in-batch… → GPT-SFT in-batch packing requires… → When CP is enabled, packed sequence… → …
  • SKILL.md covers Enablement, Code Anchors, Pitfalls and Verification
  • Calls uv

What it does

Nemo Mbridge Perf Sequence Packing is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.

Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).

It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

Example prompts

  • “/nemo-mbridge-perf-sequence-packing”

Requirements

  • Python 3

Workflow steps

12 steps, taken from the first numbered list in SKILL.md.

  1. Offline packed SFT, runtime in-batch packing, and Energon online packing are different features. Offline and Energon packing use physical…
  2. GPT-SFT in-batch packing requires dataloader_type="single" or "cyclic"; it does not support "batch".
  3. When CP is enabled, packed sequence lengths must respect 2 * context_parallel_size divisibility.
  4. For finetuning with CP, calculate_per_token_loss=True and ddp.average_in_collective=False are required.
  5. pad_cu_seqlens=True also requires pad_to_max_length=True.
  6. Packing support is model-family-specific. Qwen3-Next, GLM-4.5, and Qwen3.5-VL contain explicit opt-outs in different paths.
  7. MTP finetuning is documented as incompatible with packed sequences.
  8. Synthetic padding rows, including negative indices remapped through samples_mapping, must retain an all-zero loss mask.
  9. global_batch_size must be divisible by and no smaller than data parallel size when offline packing uses MBS1.
  10. Derive pad_seq_to_mult from CP/TP/SP for both SFT and PEFT; do not hardcode different values by workload type.
  11. pad_to_max_length controls final pack width and is conditional on fixed-shape execution requirements.
  12. Energon packing_buffer_size is per worker and also affects validation; global/eval batch counts refer to physical packs rather than source…

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemo Mbridge Perf Sequence Packing loads about 3.3k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 911 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~3.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 911 words, ~3,328 tokens.

Download SKILL.mdSave it as .claude/skills/nemo-mbridge-perf-sequence-packing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
nemo-mbridge-perf-sequence-packing
description
Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.
license
Apache-2.0

Sequence Packing Skill

For stable background and recommendation level, see:

  • @docs/training/packed-sequences.md
  • @skills/nemo-mbridge-perf-sequence-packing/card.yaml

Enablement

Offline packed SFT for LLM finetuning:

python
import math

from megatron.bridge.data.datasets.packed_sequence import PackedSequenceSpecs

cfg.train.micro_batch_size = 1
cfg.train.global_batch_size = 8
cfg.dataset.seq_length = 8192
cfg.model.seq_length = 8192
cfg.dataset.enable_offline_packing = True

cp_size = cfg.model.context_parallel_size
tp_size = cfg.model.tensor_model_parallel_size
cp_multiple = 2 * cp_size if cp_size > 1 else 1
sp_multiple = cp_size * tp_size if cfg.model.sequence_parallel and tp_size > 1 else 1
cfg.dataset.offline_packing_specs = PackedSequenceSpecs(
    packed_sequence_size=8192,
    pad_seq_to_mult=math.lcm(cp_multiple, sp_multiple),
)
Choose the offline pack length

For text-only LLM SFT and PEFT verification, start with an 8192-token offline pack when the model context limit, memory, and model-family support allow it. Benchmark pack lengths at equal token slots per optimizer step:

text
token_slots_per_step = packed_sequence_size * global_batch_size

For example, 2K/GBS32, 4K/GBS16, and 8K/GBS8 each expose 65,536 token slots per step. Longer packs aggregate more source examples into each physical MBS1 row and can reduce gradient accumulation and per-step overhead. They also increase activation memory and may expose kernel-width constraints, so select the largest measured configuration that fits rather than assuming longer is always faster.

Offline packing requires MBS1. Require global_batch_size % data_parallel_size == 0 and global_batch_size >= data_parallel_size; an 8K/GBS8 workload therefore needs DP no larger than 8. Keep model.seq_length, dataset.seq_length, and packed_sequence_size equal, use a fresh packed-data output root after changing any of them, and inspect the resolved post-setup configuration.

Equal token slots do not make different pack lengths numerically identical: the longer target changes truncation and pack membership. Rerun finite-loss, no-skip/NaN, and convergence sentinels before replacing verified evidence.

For finetuning with CP enabled:

python
cfg.model.context_parallel_size = 2
cfg.model.calculate_per_token_loss = True
cfg.ddp.average_in_collective = False

Use the same alignment formula for SFT and PEFT. It produces 1 for TP1/CP1 with SP disabled and 4 for TP4/CP1 with SP enabled. Offline packing does not derive the value automatically, so pin it explicitly and rebuild packed data after a topology change.

If a dispatcher or kernel requires a fixed final token width:

python
cfg.dataset.dataset_kwargs = {
    **(cfg.dataset.dataset_kwargs or {}),
    "pad_to_max_length": True,
}

Choose packed_sequence_size to satisfy the kernel multiple. For example, HybridEP with a 128-token combine chunk requires a width divisible by 128. This is separate from pad_seq_to_mult, which aligns each constituent sequence for CP/SP.

If CUDA graphs are enabled for this packed path, fixed token width is required and packed metadata must also have a static shape:

python
cfg.dataset.offline_packing_specs.pad_cu_seqlens = True
cfg.dataset.dataset_kwargs["pad_to_max_length"] = True

Note: pad_cu_seqlens = True also requires a metadata JSON file alongside the packed dataset (asserted in src/megatron/bridge/data/datasets/sft.py). Custom packed datasets that omit the metadata file will hit an assertion at dataset initialization.

In-batch packing for GPT SFT and supported VLM finetuning:

python
cfg.dataset.enable_in_batch_packing = True
cfg.dataset.dataloader_type = "single"
cfg.train.micro_batch_size = 4

For local or materialized GPT-SFT JSONL, this keeps the existing mmap-backed dataset and performs tokenization lazily. Both prompt/completion (GPTSFTDataset) and chat (GPTSFTChatDataset) preserve their loss-mask semantics. Use dataloader_type="single" or "cyclic" so every DataLoader yield is one logical microbatch; GPT-SFT in-batch packing does not support the global-batch "batch" dataloader.

Energon online packing for Qwen-VL uses Energon's per-worker candidate buffer instead of limiting selection to one collator micro batch:

python
cfg.dataset.packing_buffer_size = 16
cfg.dataset.micro_batch_size = 1
cfg.train.micro_batch_size = 1
cfg.model.calculate_per_token_loss = True
cfg.ddp.average_in_collective = False

packing_buffer_size is the sole native-packing selector; leave the legacy collator- and step-owned packing flags at their defaults. Use vlm_step. The buffer size counts prepared candidate samples per worker, not bytes or packed tokens. Since prepared image/video patch tensors remain in host memory until selection, start at 8-16 for high-resolution or video data and measure worker RSS, first-batch latency, and bin fill before increasing it. This path does not write offline packs; the source WebDataset shards remain unchanged. It supports eager Qwen-VL with MBS1 and rejects MTP, CUDA graphs, Qwen3-VL DistTrain, and PP. Requested MoE expert-parallel communication overlap is disabled with a warning. Standard eager alltoall EP has functional coverage for Qwen3.6-35B-A3B at TP1/PP1/EP8 with overlap disabled; this is not performance evidence. Other EP dispatchers are accepted with fixed-width native packs but do not yet have equivalent runtime evidence. The Qwen-VL model derives a MoE padding mask from logical and physical THD boundaries so fixed-width gaps do not enter auxiliary-loss, z-loss, or expert-bias statistics. Current MCore may still dispatch padded positions; expert-capacity/token-dropping configurations lack native-packing runtime coverage.

Long-context baseline:

python
cfg.model.seq_length = 16384
cfg.dataset.seq_length = 16384
cfg.model.context_parallel_size = 2
Show full SKILL.md (322 more words)Show less

Code Anchors

LLM packed SFT config surface:

128143srcmegatronbri
dataset_kwargs = {}
offline_packing_specs = None
if enable_offline_packing:
    dataset_kwargs["pad_to_max_length"] = True
    offline_packing_specs = PackedSequenceSpecs(packed_sequence_size=seq_length, pad_seq_to_mult=pad_seq_to_mult)

return _text_hf_dataset_config(
    source=HFDatasetSourceConfig(dataset_name="squad"),
    preprocessing=PromptCompletionSFTPreprocessingConfig(separator=" "),
    seq_length=seq_length,
    enable_offline_packing=enable_offline_packing,
    offline_packing_specs=offline_packing_specs,
    dataset_kwargs=dataset_kwargs,
    val_proportion=0.1,
    num_workers=1,
)

The shared text-dataset helper currently opts into fixed-width packs. Treat that as a helper default, not a universal offline-packing runtime requirement; preserve it when the selected dispatcher, kernel, or CUDA-graph path requires static width.

Bridge validation:

12201248srcmegatronb
enable_in_batch_packing = getattr(self.dataset, "enable_in_batch_packing", False)
enable_offline_packing = getattr(self.dataset, "enable_offline_packing", False)
offline_packing_specs = getattr(self.dataset, "offline_packing_specs", None)

if enable_offline_packing and enable_in_batch_packing:
    raise ValueError("enable_offline_packing and enable_in_batch_packing are mutually exclusive.")
if enable_offline_packing and offline_packing_specs is None:
    raise ValueError("offline_packing_specs must be set when enable_offline_packing=True.")
...
if enable_in_batch_packing:
    ...
    cp_multiple = 2 * cp_size if cp_size > 1 else 1
    sp_multiple = cp_size * tp_size if has_sp and tp_size > 1 else 1
    self.dataset.in_batch_packing_pad_to_multiple_of = math.lcm(cp_multiple, sp_multiple)
14001442srcmegatronb
if self.model.context_parallel_size > 1:
    assert self.model.seq_length % (self.model.context_parallel_size * 2) == 0, ...
    if isinstance(self.dataset, FinetuningDatasetConfig):
        assert self.model.calculate_per_token_loss, ...
        assert not self.ddp.average_in_collective, ...
...
if enable_offline_packing and self.train.micro_batch_size > 1:
    raise ValueError(...)
...
if enable_in_batch_packing and self.train.micro_batch_size == 1:
    raise ValueError(...)

Collate-time in-batch runtime used by VLM providers:

397449srcmegatronbri
def prepare_padded_or_packed_sequence_batch(
    batch,
    *,
    sequence_length,
    ...
    enable_in_batch_packing=False,
    in_batch_packing_pad_to_multiple_of=1,
    ...
):
    ...
    if enable_in_batch_packing:
        pack_right_padded_sequence_batch_to_mcore_thd(
            batch,
            sequence_length=sequence_length,
            pad_to_multiple_of=in_batch_packing_pad_to_multiple_of,
            ...
        )
        return

GPT-SFT direct-row packing:

627671srcmegatronbri
def _collate_in_batch(self, batch):
    ...
    return build_mcore_thd_sequence_batch_from_rows(...)

Packed THD runtime constraint:

94108srcmegatronbrid
if batch.get("cu_seqlens_q") is not None:
    cu_seqlens = batch.get("cu_seqlens_q_padded")
    if cu_seqlens is None:
        cu_seqlens = batch["cu_seqlens_q"]
    if cu_seqlens.dim() > 1 and cu_seqlens.size(0) != 1:
        raise ValueError("Packed THD batches expect micro-batch size 1 for context-parallel slicing (THD layout)")
    return cu_seqlens.squeeze()

cu_seqlens = batch["cu_seqlens"]
if cu_seqlens.dim() > 1 and cu_seqlens.size(0) != 1:
    raise ValueError("Packed THD batches expect micro-batch size 1 for context-parallel slicing (THD layout)")

Pitfalls

  1. Offline packed SFT, runtime in-batch packing, and Energon online packing are different features. Offline and Energon packing use physical MBS1; runtime in-batch packing uses MBS greater than one.
  2. GPT-SFT in-batch packing requires dataloader_type="single" or "cyclic"; it does not support "batch".
  3. When CP is enabled, packed sequence lengths must respect 2 * context_parallel_size divisibility.
  4. For finetuning with CP, calculate_per_token_loss=True and ddp.average_in_collective=False are required.
  5. pad_cu_seqlens=True also requires pad_to_max_length=True.
  6. Packing support is model-family-specific. Qwen3-Next, GLM-4.5, and Qwen3.5-VL contain explicit opt-outs in different paths.
  7. MTP finetuning is documented as incompatible with packed sequences.
  8. Synthetic padding rows, including negative indices remapped through samples_mapping, must retain an all-zero loss mask.
  9. global_batch_size must be divisible by and no smaller than data parallel size when offline packing uses MBS1.
  10. Derive pad_seq_to_mult from CP/TP/SP for both SFT and PEFT; do not hardcode different values by workload type.
  11. pad_to_max_length controls final pack width and is conditional on fixed-shape execution requirements.
  12. Energon packing_buffer_size is per worker and also affects validation; global/eval batch counts refer to physical packs rather than source conversations.
  13. Exact Energon loader resume requires unchanged shards/splits, DP world size, worker counts, shuffle settings/seed, processor, sequence length, topology, and packing-buffer size.

Verification

Use the checked-in unit coverage:

bash
uv run python -m pytest tests/unit_tests/training/utils/test_packed_seq_utils.py -v && \
uv run python -m pytest tests/unit_tests/training/test_config.py -k "packed_sequence or enable_in_batch_packing or offline_and_in_batch_packing_are_mutually_exclusive or context_parallel_seq_length_divisibility or context_parallel_finetuning_validations" -v && \
uv run python -m pytest tests/unit_tests/data/packing/test_in_batch.py -v && \
uv run python -m pytest tests/unit_tests/data/datasets/test_gpt_sft.py -k "in_batch_packing" -v && \
uv run python -m pytest tests/unit_tests/data/builders/test_gpt_sft_config.py -v && \
uv run python -m pytest tests/unit_tests/training/test_vlm_step.py -k "deferred_in_batch_packing or packed_metadata" -v && \
uv run python -m pytest tests/unit_tests/models/qwen_vl/data/test_energon.py tests/unit_tests/data/builders/test_energon_builder.py -v && \
uv run python -m pytest tests/unit_tests/tutorials/test_multimodal_data_tutorials.py -k "native_packing_loader" -v && \
uv run python -m pytest tests/unit_tests/data/datasets/test_packed_parquet.py -k "negative_index_zeroes_loss_mask" -v && \
uv run python -m pytest tests/unit_tests/data/datasets/test_sft.py -k "mapped_padding_rows_do_not_contribute_to_loss" -v

Success criteria:

  • all selected tests pass
  • offline and in-batch configuration validation remains mutually exclusive
  • packed metadata reaches the training step in MCore THD form
  • GPT-SFT in-batch packing rejects the global-batch "batch" dataloader
  • native Energon packing restores pending groups exactly and flushes finite partial buffers without dropping samples
  • mapped padding rows do not contribute to loss

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in skills/nemo-mbridge-perf-sequence-packing of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • card.yaml
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Nemo Mbridge Perf Sequence Packing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemo Mbridge Perf Sequence Packing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemo Mbridge Perf Sequence Packing this skillNVIDIA/skills3.5k—~3.3kAutomated safety check: PassApache-2.0
Skill InspectorNVIDIA/SkillSpector20k1 repos~1.8kAutomated safety check: PassApache-2.0
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
NEAR AI Cloud Private Inferenceinternet-court/internet-court-skill6.4k2 repos~1.3kAutomated safety check: PassCustom licence
Nemoclaw Maintainer Normalize Title TagsNVIDIA/NemoClaw23k—~693Automated safety check: PassApache-2.0

Similar skills

  • Skill Inspector

    NVIDIA/SkillSpector

    Official

    Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.

    20k GitHub starsUsed in 1 repo~1.8k tokens
    SecurityAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • NEAR AI Cloud Private Inference

    internet-court/internet-court-skill

    Shows how to call NEAR AI Cloud through an OpenAI-compatible API and verify that inference ran in a TEE, using attestation checks and signed chat responses.

    6.4k GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Remove bracketed NemoClaw tags from GitHub issue and PR titles.

    23k GitHub stars~693 tokensUpdated today
    Marketing & SEOAuto-check passed
  • Audit and implement a NemoClaw dependency version upgrade, including Hermes and base images.

    23k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemo Mbridge Perf Sequence Packing

What does Nemo Mbridge Perf Sequence Packing do?

Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints. Nemo Mbridge Perf Sequence Packing is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.

How do I install Nemo Mbridge Perf Sequence Packing in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-sequence-packing in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-sequence-packing in your project. Claude Code loads it when a task matches its description.

How do I install Nemo Mbridge Perf Sequence Packing in Codex?

Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-sequence-packing in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-sequence-packing in your project. Codex loads it when a task matches its description.

Can I use Nemo Mbridge Perf Sequence Packing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-sequence-packing, .gemini/skills/nemo-mbridge-perf-sequence-packing, .github/skills/nemo-mbridge-perf-sequence-packing and .opencode/skills/nemo-mbridge-perf-sequence-packing in your project.

What does Nemo Mbridge Perf Sequence Packing need to run?

Going by SKILL.md and its folder, Nemo Mbridge Perf Sequence Packing needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Nemo Mbridge Perf Sequence Packing access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Nemo Mbridge Perf Sequence Packing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemo Mbridge Perf Sequence Packing use?

Nemo Mbridge Perf Sequence Packing is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemo Mbridge Perf Sequence Packing use?

About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Nemo Mbridge Perf Sequence Packing?

Skills that share tags, products or a category with Nemo Mbridge Perf Sequence Packing: Skill Inspector (NVIDIA/SkillSpector, 20k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Embeddings via 9Router (decolua/9router, 30k stars) and NEAR AI Cloud Private Inference (internet-court/internet-court-skill, 6.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemo Mbridge Perf Sequence Packing?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.