Tilelang Skill
slowlyC/agent-gpu-skills
Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.
Evidence-gated workflow for MoE performance optimization in Megatron Bridge.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-optimization-workflow --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-optimization-workflow .claude/skills/nemo-mbridge-perf-moe-optimization-workflow && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-moe-optimization-workflow" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-optimization-workflow into .claude/skills/nemo-mbridge-perf-moe-optimization-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-optimization-workflow", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-optimization-workflowType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-optimization-workflow --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-optimization-workflow .agents/skills/nemo-mbridge-perf-moe-optimization-workflow && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-moe-optimization-workflow" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-optimization-workflow into .agents/skills/nemo-mbridge-perf-moe-optimization-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-optimization-workflow", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-optimization-workflow --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-optimization-workflow .cursor/skills/nemo-mbridge-perf-moe-optimization-workflow && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-moe-optimization-workflow" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-optimization-workflow into .cursor/skills/nemo-mbridge-perf-moe-optimization-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-optimization-workflow", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-moe-optimization-workflow--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-optimization-workflow --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-optimization-workflow .gemini/skills/nemo-mbridge-perf-moe-optimization-workflow && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-optimization-workflow" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-optimization-workflow into .gemini/skills/nemo-mbridge-perf-moe-optimization-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-optimization-workflow", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-optimization-workflowInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-optimization-workflow .github/skills/nemo-mbridge-perf-moe-optimization-workflow && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-optimization-workflow" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-optimization-workflow into .github/skills/nemo-mbridge-perf-moe-optimization-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-optimization-workflow", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-optimization-workflow --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-optimization-workflow .opencode/skills/nemo-mbridge-perf-moe-optimization-workflow && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-optimization-workflow" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-optimization-workflow into .opencode/skills/nemo-mbridge-perf-moe-optimization-workflow/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-optimization-workflow", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-moe-optimization-workflowEvidence-gated workflow for MoE performance optimization in Megatron Bridge.
Nemo Mbridge Perf Moe Optimization Workflow is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Evidence-gated workflow for MoE performance optimization in Megatron Bridge. Covers measurement contracts, the Three Walls framework, parallel folding, profiling, matched A/B tuning, and final validation.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It sits in Development, covering Performance optimization. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
arxiv.orgFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Moe Optimization Workflow loads about 3.3k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 1,656 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,656 words, ~3,259 tokens.
.claude/skills/nemo-mbridge-perf-moe-optimization-workflow/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-optimization-workflow/card.yaml Source: Scalable Training of MoE Models with Megatron Core
Start with the paper's Three Walls:
For operational diagnosis, split the compute-efficiency wall into compute and host/launch bottlenecks. They need different evidence and different fixes. MoE tuning is iterative, so use this order:
freeze the measurement contract -> fit -> scale -> profile -> retune -> validateFor MoE optimization workflow prompts, present the response in this order:
--fake-init-process-group to sanity-check large layouts.Attention: TP x CP x DP x PP
and MoE: ETP x EP x EDP x PP.alltoall for safe bring-up, then A/B flex + deepep and flex +
hybridep when their packages and target topology support them. Start from
BF16 and eager execution; introduce lower precision or the narrowest useful
CUDA-graph scope only after profiling justifies it.A comparison is valid only when the following stay fixed unless they are the single variable under test:
Separate two acceptance classes:
Start with a configuration that fits reliably before chasing throughput.
Recommended order:
--fake-init-process-group to sanity-check large parallel layouts on a
single GPU before burning cluster time.Prefer selective recompute for MoE runs:
layernorm, core_attn, moe_act, mlp, or
model-specific modules (shared_experts, mla_up_proj)As a rule of thumb, fine-grained recompute often recovers most of the needed memory while keeping throughput much closer to the non-recompute baseline than full-layer recompute does.
Priority order:
Parallel Folding decouples attention and MoE parallelism so you do not have to pick a single compromise layout:
Attention: TP × CP × DP × PP
MoE: ETP × EP × EDP × PPKey knobs:
--expert-model-parallel-size--expert-tensor-parallel-sizeUse it when attention prefers some TP or CP, but expert layers benefit from a larger EP degree than the dense layers can tolerate.
| Bottleneck | What it looks like | Primary fixes |
|---|---|---|
| Memory | Run fits only with aggressive full recompute or OOMs during warmup | selective recompute, FP8, offloading, better PP layout |
| Communication | Nsight shows large all-to-all or collective blocks | DeepEP or HybridEP, EP overlap, DP/TP overlap, better PP layout |
| Host overhead | GPU gaps, launch-bound traces, Python overhead | CUDA graphs, --manual-gc, higher MBS, CPU affinity tuning |
| Compute | Low SM utilization after comm and host issues are addressed | grouped GEMM, fusion work, FP8, dispatcher-specific kernel tuning |
Use unprofiled steady iterations for the acceptance metric and a matched profile for causal explanation:
On a controlled 16×H100 Qwen3 30B-A3B HybridEP run, plain EP overlap increased
communication hidden by GEMM/attention from 0.11% to 36.55%. The unprofiled
step fell from 24.7138s to 20.9920s and throughput rose from 244.039 to 287.305
model TFLOPS/GPU. delay_wgrad_compute remained disabled.
Choose the smallest candidate that targets the profiled bottleneck and change one variable at a time.
Use dispatcher choice as a bottleneck fix, not as a hardware lookup table.
moe_token_dispatcher_type="alltoall": safest bring-up path, fine for
smaller EP sizesmoe_token_dispatcher_type="flex" + moe_flex_dispatcher_backend="deepep":
candidate when DeepEP is installed and communication is exposedmoe_token_dispatcher_type="flex" + moe_flex_dispatcher_backend="hybridep":
topology-sensitive candidate on both NVL8 and NVL72 systems when HybridEP is
installedHybridEP plus plain EP overlap is the current measured winner for the canonical
16×H100 Qwen3 30B-A3B shape, while the canonical 256×H100 Qwen3 235B recipe
uses standard alltoall plus overlap. Benchmark backend compatibility and
throughput in the target container; neither GPU name nor EP degree determines
the winner by itself.
If the all-to-all path is visible in profiles, combine dispatcher tuning with:
--overlap-moe-expert-parallel-comm--overlap-grad-reduce--tp-comm-overlapTest plain EP overlap, shared-expert overlap, and delayed weight-gradient compute as separate candidates first. A combination can regress even when one component helped on another model.
Start with a verified BF16 baseline. Hardware capability only determines which lower-precision candidates are legal; it does not guarantee a speedup.
| Platform | Candidate after BF16 is stable |
|---|---|
| Hopper | per-tensor, current-scaling, or blockwise FP8 supported by the target stack |
| Blackwell | MXFP8 or another supported FP8 recipe |
| Blackwell, speed-first exploration | NVFP4 after the BF16/FP8 path is stable |
Keep the router in FP32. The largest wins usually come from expert GEMMs and other heavy matrix math, not from trying to quantize every small MoE component. Require logs or traces showing that the intended kernels ran, and judge the candidate by end-to-end steady step time rather than theoretical peak FLOPS.
Use CUDA graphs only after a profile shows meaningful host/launch gaps. For dropless MoE, start with the narrowest partial TE-scoped graph candidate:
moe_routermoe_preprocessAdd attn only if it is supported for the model and improves the same matched
stack. A successful capture is not evidence of a speedup, and a graph win can
disappear after dispatcher, overlap, or precision changes.
This path keeps dynamic expert work outside the graph. Budget extra memory, verify that shapes remain static, confirm replay rather than capture alone, and time only post-capture iterations.
Use full-iteration graphs only for graph-friendly workloads such as drop-and-pad or tightly controlled static-shape experiments.
Related references:
Use 6–12 post-warmup iterations for inexpensive screening when the workload allows it. For the selected candidate, run at least 50 steps and report a fixed steady window such as steps 41–50. The final evidence bundle should contain:
Do not attribute the total gain of a final multi-change winner to one earlier A/B. For example, the Qwen3 overlap experiment isolated a rise from 244.039 to 287.305 TFLOPS/GPU; the later canonical recipe reached 299.352 after additional HybridEP tuning. They answer different questions.
Do not optimize in the wrong order: fitting the model and selecting sane parallelism matter more than micro-optimizations.
Platform changes the limiting wall: H100-class runs often feel more communication-bound, while GB200 or GB300 runs often expose CPU or launch overhead earlier.
FP8 MFU can look misleadingly low: compare absolute throughput as well as MFU when switching precision modes.
CUDA graphs and recompute interact: TE-scoped graphs are usually paired with selective recompute, not blanket full recompute.
Parallel Folding is not optional at large scale: once attention and expert layers want clearly different layouts, a single shared TP or EP plan becomes a tax on both.
Summed kernel time is not exposed time: use interval unions and communication/compute intersection when validating overlap.
Benchmark-only semantics are not production acceptance: forced routing, synthetic data, or disabled optimizer/checkpoint paths must be disclosed and validated separately from training-equivalent results.
Feature activation needs evidence: a config dump is insufficient when a backend can fall back, a graph can capture without helping, or a lower- precision recipe can miss the intended kernels.
Last signature refresh: 2026-08-03.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-optimization-workflow of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Nemo Mbridge Perf Moe Optimization Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Moe Optimization Workflow this skillNVIDIA/skills | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Tilelang SkillslowlyC/agent-gpu-skills | 169 | — | ~1.8k | Automated safety check: Pass | MIT | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~2.8k | Automated safety check: Pass | None | |
| Cuda Profilingmohitmishra786/low-level-dev-skills | 253 | — | ~1.6k | Automated safety check: Notes | MIT | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Analysissgl-project/sglang | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 |
slowlyC/agent-gpu-skills
Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
mohitmishra786/low-level-dev-skills
CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Evidence-gated workflow for MoE performance optimization in Megatron Bridge. Nemo Mbridge Perf Moe Optimization Workflow is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Evidence-gated workflow for MoE performance optimization in Megatron Bridge.
Nemo Mbridge Perf Moe Optimization Workflow fits situations like: tasks that involve Performance optimization.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-optimization-workflow in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-optimization-workflow in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-optimization-workflow in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-optimization-workflow in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-optimization-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-optimization-workflow, .gemini/skills/nemo-mbridge-perf-moe-optimization-workflow, .github/skills/nemo-mbridge-perf-moe-optimization-workflow and .opencode/skills/nemo-mbridge-perf-moe-optimization-workflow in your project.
SKILL.md names no scripts, command-line tools or credentials: Nemo Mbridge Perf Moe Optimization Workflow is instructions for the agent only. Our summary lists: Python 3.
SKILL.md names 1 domain. As links in the text: arxiv.org. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Moe Optimization Workflow is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Moe Optimization Workflow: Tilelang Skill (slowlyC/agent-gpu-skills, 169 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Cuda Profiling (mohitmishra786/low-level-dev-skills, 253 stars) and Graphsignal (graphsignal/graphsignal, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.