Skill Inspector
NVIDIA/SkillSpector
Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.
Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-sequence-packing --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-sequence-packing .claude/skills/nemo-mbridge-perf-sequence-packing && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-sequence-packing" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-sequence-packing into .claude/skills/nemo-mbridge-perf-sequence-packing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-sequence-packing", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-sequence-packingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-sequence-packing --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-sequence-packing .agents/skills/nemo-mbridge-perf-sequence-packing && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-sequence-packing" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-sequence-packing into .agents/skills/nemo-mbridge-perf-sequence-packing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-sequence-packing", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-sequence-packing --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-sequence-packing .cursor/skills/nemo-mbridge-perf-sequence-packing && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-sequence-packing" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-sequence-packing into .cursor/skills/nemo-mbridge-perf-sequence-packing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-sequence-packing", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-sequence-packing--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-sequence-packing --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-sequence-packing .gemini/skills/nemo-mbridge-perf-sequence-packing && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-sequence-packing" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-sequence-packing into .gemini/skills/nemo-mbridge-perf-sequence-packing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-sequence-packing", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-sequence-packingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-sequence-packing .github/skills/nemo-mbridge-perf-sequence-packing && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-sequence-packing" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-sequence-packing into .github/skills/nemo-mbridge-perf-sequence-packing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-sequence-packing", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-sequence-packing --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-sequence-packing .opencode/skills/nemo-mbridge-perf-sequence-packing && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-sequence-packing" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-sequence-packing into .opencode/skills/nemo-mbridge-perf-sequence-packing/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-sequence-packing", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-sequence-packingValidate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.
Nemo Mbridge Perf Sequence Packing is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
12 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Sequence Packing loads about 3.3k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 911 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 911 words, ~3,328 tokens.
.claude/skills/nemo-mbridge-perf-sequence-packing/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.For stable background and recommendation level, see:
Offline packed SFT for LLM finetuning:
import math
from megatron.bridge.data.datasets.packed_sequence import PackedSequenceSpecs
cfg.train.micro_batch_size = 1
cfg.train.global_batch_size = 8
cfg.dataset.seq_length = 8192
cfg.model.seq_length = 8192
cfg.dataset.enable_offline_packing = True
cp_size = cfg.model.context_parallel_size
tp_size = cfg.model.tensor_model_parallel_size
cp_multiple = 2 * cp_size if cp_size > 1 else 1
sp_multiple = cp_size * tp_size if cfg.model.sequence_parallel and tp_size > 1 else 1
cfg.dataset.offline_packing_specs = PackedSequenceSpecs(
packed_sequence_size=8192,
pad_seq_to_mult=math.lcm(cp_multiple, sp_multiple),
)For text-only LLM SFT and PEFT verification, start with an 8192-token offline pack when the model context limit, memory, and model-family support allow it. Benchmark pack lengths at equal token slots per optimizer step:
token_slots_per_step = packed_sequence_size * global_batch_sizeFor example, 2K/GBS32, 4K/GBS16, and 8K/GBS8 each expose 65,536 token slots per step. Longer packs aggregate more source examples into each physical MBS1 row and can reduce gradient accumulation and per-step overhead. They also increase activation memory and may expose kernel-width constraints, so select the largest measured configuration that fits rather than assuming longer is always faster.
Offline packing requires MBS1. Require global_batch_size % data_parallel_size == 0 and global_batch_size >= data_parallel_size; an 8K/GBS8 workload
therefore needs DP no larger than 8. Keep model.seq_length,
dataset.seq_length, and packed_sequence_size equal, use a fresh packed-data
output root after changing any of them, and inspect the resolved post-setup
configuration.
Equal token slots do not make different pack lengths numerically identical: the longer target changes truncation and pack membership. Rerun finite-loss, no-skip/NaN, and convergence sentinels before replacing verified evidence.
For finetuning with CP enabled:
cfg.model.context_parallel_size = 2
cfg.model.calculate_per_token_loss = True
cfg.ddp.average_in_collective = FalseUse the same alignment formula for SFT and PEFT. It produces 1 for TP1/CP1 with SP disabled and 4 for TP4/CP1 with SP enabled. Offline packing does not derive the value automatically, so pin it explicitly and rebuild packed data after a topology change.
If a dispatcher or kernel requires a fixed final token width:
cfg.dataset.dataset_kwargs = {
**(cfg.dataset.dataset_kwargs or {}),
"pad_to_max_length": True,
}Choose packed_sequence_size to satisfy the kernel multiple. For example,
HybridEP with a 128-token combine chunk requires a width divisible by 128.
This is separate from pad_seq_to_mult, which aligns each constituent
sequence for CP/SP.
If CUDA graphs are enabled for this packed path, fixed token width is required and packed metadata must also have a static shape:
cfg.dataset.offline_packing_specs.pad_cu_seqlens = True
cfg.dataset.dataset_kwargs["pad_to_max_length"] = TrueNote: pad_cu_seqlens = True also requires a metadata JSON file alongside
the packed dataset (asserted in src/megatron/bridge/data/datasets/sft.py).
Custom packed datasets that omit the metadata file will hit an assertion at
dataset initialization.
In-batch packing for GPT SFT and supported VLM finetuning:
cfg.dataset.enable_in_batch_packing = True
cfg.dataset.dataloader_type = "single"
cfg.train.micro_batch_size = 4For local or materialized GPT-SFT JSONL, this keeps the existing mmap-backed
dataset and performs tokenization lazily. Both prompt/completion
(GPTSFTDataset) and chat (GPTSFTChatDataset) preserve their loss-mask
semantics. Use dataloader_type="single" or "cyclic" so every DataLoader
yield is one logical microbatch; GPT-SFT in-batch packing does not support the
global-batch "batch" dataloader.
Energon online packing for Qwen-VL uses Energon's per-worker candidate buffer instead of limiting selection to one collator micro batch:
cfg.dataset.packing_buffer_size = 16
cfg.dataset.micro_batch_size = 1
cfg.train.micro_batch_size = 1
cfg.model.calculate_per_token_loss = True
cfg.ddp.average_in_collective = Falsepacking_buffer_size is the sole native-packing selector; leave the legacy
collator- and step-owned packing flags at their defaults. Use vlm_step. The
buffer size counts prepared candidate samples per worker, not bytes or packed
tokens. Since prepared image/video patch tensors remain in
host memory until selection, start at 8-16 for high-resolution or video data
and measure worker RSS, first-batch latency, and bin fill before increasing it.
This path does not write offline packs; the source WebDataset shards remain
unchanged. It supports eager Qwen-VL with MBS1 and rejects MTP, CUDA graphs,
Qwen3-VL DistTrain, and PP. Requested MoE expert-parallel communication overlap
is disabled with a warning. Standard eager alltoall EP has functional coverage
for Qwen3.6-35B-A3B at TP1/PP1/EP8 with overlap disabled; this is not performance
evidence. Other EP dispatchers are accepted with fixed-width native packs but do
not yet have equivalent runtime evidence. The Qwen-VL model derives a MoE
padding mask from logical and physical THD boundaries so fixed-width gaps do not
enter auxiliary-loss, z-loss, or expert-bias statistics. Current MCore may still
dispatch padded positions; expert-capacity/token-dropping configurations lack
native-packing runtime coverage.
Long-context baseline:
cfg.model.seq_length = 16384
cfg.dataset.seq_length = 16384
cfg.model.context_parallel_size = 2LLM packed SFT config surface:
dataset_kwargs = {}
offline_packing_specs = None
if enable_offline_packing:
dataset_kwargs["pad_to_max_length"] = True
offline_packing_specs = PackedSequenceSpecs(packed_sequence_size=seq_length, pad_seq_to_mult=pad_seq_to_mult)
return _text_hf_dataset_config(
source=HFDatasetSourceConfig(dataset_name="squad"),
preprocessing=PromptCompletionSFTPreprocessingConfig(separator=" "),
seq_length=seq_length,
enable_offline_packing=enable_offline_packing,
offline_packing_specs=offline_packing_specs,
dataset_kwargs=dataset_kwargs,
val_proportion=0.1,
num_workers=1,
)The shared text-dataset helper currently opts into fixed-width packs. Treat that as a helper default, not a universal offline-packing runtime requirement; preserve it when the selected dispatcher, kernel, or CUDA-graph path requires static width.
Bridge validation:
enable_in_batch_packing = getattr(self.dataset, "enable_in_batch_packing", False)
enable_offline_packing = getattr(self.dataset, "enable_offline_packing", False)
offline_packing_specs = getattr(self.dataset, "offline_packing_specs", None)
if enable_offline_packing and enable_in_batch_packing:
raise ValueError("enable_offline_packing and enable_in_batch_packing are mutually exclusive.")
if enable_offline_packing and offline_packing_specs is None:
raise ValueError("offline_packing_specs must be set when enable_offline_packing=True.")
...
if enable_in_batch_packing:
...
cp_multiple = 2 * cp_size if cp_size > 1 else 1
sp_multiple = cp_size * tp_size if has_sp and tp_size > 1 else 1
self.dataset.in_batch_packing_pad_to_multiple_of = math.lcm(cp_multiple, sp_multiple)if self.model.context_parallel_size > 1:
assert self.model.seq_length % (self.model.context_parallel_size * 2) == 0, ...
if isinstance(self.dataset, FinetuningDatasetConfig):
assert self.model.calculate_per_token_loss, ...
assert not self.ddp.average_in_collective, ...
...
if enable_offline_packing and self.train.micro_batch_size > 1:
raise ValueError(...)
...
if enable_in_batch_packing and self.train.micro_batch_size == 1:
raise ValueError(...)Collate-time in-batch runtime used by VLM providers:
def prepare_padded_or_packed_sequence_batch(
batch,
*,
sequence_length,
...
enable_in_batch_packing=False,
in_batch_packing_pad_to_multiple_of=1,
...
):
...
if enable_in_batch_packing:
pack_right_padded_sequence_batch_to_mcore_thd(
batch,
sequence_length=sequence_length,
pad_to_multiple_of=in_batch_packing_pad_to_multiple_of,
...
)
returnGPT-SFT direct-row packing:
def _collate_in_batch(self, batch):
...
return build_mcore_thd_sequence_batch_from_rows(...)Packed THD runtime constraint:
if batch.get("cu_seqlens_q") is not None:
cu_seqlens = batch.get("cu_seqlens_q_padded")
if cu_seqlens is None:
cu_seqlens = batch["cu_seqlens_q"]
if cu_seqlens.dim() > 1 and cu_seqlens.size(0) != 1:
raise ValueError("Packed THD batches expect micro-batch size 1 for context-parallel slicing (THD layout)")
return cu_seqlens.squeeze()
cu_seqlens = batch["cu_seqlens"]
if cu_seqlens.dim() > 1 and cu_seqlens.size(0) != 1:
raise ValueError("Packed THD batches expect micro-batch size 1 for context-parallel slicing (THD layout)")dataloader_type="single" or "cyclic"; it does not support "batch".2 * context_parallel_size divisibility.calculate_per_token_loss=True and ddp.average_in_collective=False are required.pad_cu_seqlens=True also requires pad_to_max_length=True.Qwen3-Next, GLM-4.5, and Qwen3.5-VL contain explicit opt-outs in different paths.samples_mapping, must retain an all-zero loss mask.global_batch_size must be divisible by and no smaller than data parallel size when offline packing uses MBS1.pad_seq_to_mult from CP/TP/SP for both SFT and PEFT; do not hardcode different values by workload type.pad_to_max_length controls final pack width and is conditional on fixed-shape execution requirements.packing_buffer_size is per worker and also affects validation; global/eval batch counts refer to physical packs rather than source conversations.Use the checked-in unit coverage:
uv run python -m pytest tests/unit_tests/training/utils/test_packed_seq_utils.py -v && \
uv run python -m pytest tests/unit_tests/training/test_config.py -k "packed_sequence or enable_in_batch_packing or offline_and_in_batch_packing_are_mutually_exclusive or context_parallel_seq_length_divisibility or context_parallel_finetuning_validations" -v && \
uv run python -m pytest tests/unit_tests/data/packing/test_in_batch.py -v && \
uv run python -m pytest tests/unit_tests/data/datasets/test_gpt_sft.py -k "in_batch_packing" -v && \
uv run python -m pytest tests/unit_tests/data/builders/test_gpt_sft_config.py -v && \
uv run python -m pytest tests/unit_tests/training/test_vlm_step.py -k "deferred_in_batch_packing or packed_metadata" -v && \
uv run python -m pytest tests/unit_tests/models/qwen_vl/data/test_energon.py tests/unit_tests/data/builders/test_energon_builder.py -v && \
uv run python -m pytest tests/unit_tests/tutorials/test_multimodal_data_tutorials.py -k "native_packing_loader" -v && \
uv run python -m pytest tests/unit_tests/data/datasets/test_packed_parquet.py -k "negative_index_zeroes_loss_mask" -v && \
uv run python -m pytest tests/unit_tests/data/datasets/test_sft.py -k "mapped_padding_rows_do_not_contribute_to_loss" -vSuccess criteria:
"batch" dataloader© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-sequence-packing of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Nemo Mbridge Perf Sequence Packing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Sequence Packing this skillNVIDIA/skills | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Skill InspectorNVIDIA/SkillSpector | 20k | 1 repos | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Analysissgl-project/sglang | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| Embeddings via 9Routerdecolua/9router | 30k | — | ~604 | Automated safety check: Pass | MIT | |
| NEAR AI Cloud Private Inferenceinternet-court/internet-court-skill | 6.4k | 2 repos | ~1.3k | Automated safety check: Pass | Custom licence | |
| Nemoclaw Maintainer Normalize Title TagsNVIDIA/NemoClaw | 23k | — | ~693 | Automated safety check: Pass | Apache-2.0 |
NVIDIA/SkillSpector
Decides whether an agent skill is safe to install by combining a SkillSpector static scan with the agent's own source review, ending in APPROVE, CAUTION or REJECT.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
decolua/9router
Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.
internet-court/internet-court-skill
Shows how to call NEAR AI Cloud through an OpenAI-compatible API and verify that inference ran in a TEE, using attestation checks and signed chat responses.
NVIDIA/NemoClaw
Remove bracketed NemoClaw tags from GitHub issue and PR titles.
NVIDIA/NemoClaw
Audit and implement a NemoClaw dependency version upgrade, including Hermes and base images.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints. Nemo Mbridge Perf Sequence Packing is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Validate and use packed sequences and long-context training in Megatron-Bridge, including offline LLM packing, collate-time VLM packing, Energon online packing, and CP constraints.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-sequence-packing in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-sequence-packing in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-sequence-packing in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-sequence-packing in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-sequence-packing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-sequence-packing, .gemini/skills/nemo-mbridge-perf-sequence-packing, .github/skills/nemo-mbridge-perf-sequence-packing and .opencode/skills/nemo-mbridge-perf-sequence-packing in your project.
Going by SKILL.md and its folder, Nemo Mbridge Perf Sequence Packing needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Sequence Packing is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Sequence Packing: Skill Inspector (NVIDIA/SkillSpector, 20k stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Embeddings via 9Router (decolua/9router, 30k stars) and NEAR AI Cloud Private Inference (internet-court/internet-court-skill, 6.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.