Add New Model
JakeATX/llamAmpere
Guided workflow for adding a new model architecture to llama.cpp.
Representative, point-in-time MoE training playbooks by hardware and model family.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-hardware-configs --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-hardware-configs .claude/skills/nemo-mbridge-perf-moe-hardware-configs && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemo-mbridge-perf-moe-hardware-configs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-hardware-configs into .claude/skills/nemo-mbridge-perf-moe-hardware-configs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-hardware-configs", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-hardware-configsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-hardware-configs --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-hardware-configs .agents/skills/nemo-mbridge-perf-moe-hardware-configs && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemo-mbridge-perf-moe-hardware-configs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-hardware-configs into .agents/skills/nemo-mbridge-perf-moe-hardware-configs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-hardware-configs", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-hardware-configs --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-hardware-configs .cursor/skills/nemo-mbridge-perf-moe-hardware-configs && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemo-mbridge-perf-moe-hardware-configs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-hardware-configs into .cursor/skills/nemo-mbridge-perf-moe-hardware-configs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-hardware-configs", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemo-mbridge-perf-moe-hardware-configs--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-hardware-configs --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-hardware-configs .gemini/skills/nemo-mbridge-perf-moe-hardware-configs && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-hardware-configs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-hardware-configs into .gemini/skills/nemo-mbridge-perf-moe-hardware-configs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-hardware-configs", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-hardware-configsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-hardware-configs .github/skills/nemo-mbridge-perf-moe-hardware-configs && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-hardware-configs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-hardware-configs into .github/skills/nemo-mbridge-perf-moe-hardware-configs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-hardware-configs", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemo-mbridge-perf-moe-hardware-configs --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemo-mbridge-perf-moe-hardware-configs .opencode/skills/nemo-mbridge-perf-moe-hardware-configs && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemo-mbridge-perf-moe-hardware-configs" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemo-mbridge-perf-moe-hardware-configs into .opencode/skills/nemo-mbridge-perf-moe-hardware-configs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemo-mbridge-perf-moe-hardware-configs", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemo-mbridge-perf-moe-hardware-configsRepresentative, point-in-time MoE training playbooks by hardware and model family.
Nemo Mbridge Perf Moe Hardware Configs is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Representative, point-in-time MoE training playbooks by hardware and model family. Use them as candidate seeds, then revalidate the exact runtime, semantics, topology, and steady-state throughput.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `BENCHMARK.md`, `card.yaml` and `evals/evals.json`).
It sits in AI & LLM Engineering. It works with CUDA and Qwen. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 14a98ae. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemo Mbridge Perf Moe Hardware Configs loads about 2k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 722 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 14a98ae, republished under its Apache-2.0 licence (© NVIDIA). 722 words, ~1,984 tokens.
.claude/skills/nemo-mbridge-perf-moe-hardware-configs/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Stable docs: @docs/training/moe-optimization.md Card: @skills/nemo-mbridge-perf-moe-hardware-configs/card.yaml
These rows are search seeds, not hardware defaults or throughput promises.
| Platform | Candidates to screen after alltoall bring-up | What usually matters most |
|---|---|---|
| H100 | DeepEP or HybridEP, explicit overlap, supported FP8 modes | communication overlap, dispatcher/runtime compatibility, and PP efficiency |
| B200 | DeepEP or HybridEP, supported FP8 modes, careful PP layout | container quality and tuned communication settings |
| GB200 | HybridEP, then profile-driven graphs and CPU cleanup | host overhead, topology-aware dispatch, memory headroom |
| GB300 | HybridEP and the target container's lower-precision/kernel stack | the same system interactions as GB200, with remeasurement required |
For hardware playbook questions, answer from these canonical rows before adding throughput caveats:
| Workload | Hardware | Dispatcher | Layout |
|---|---|---|---|
| DSV3 | H100 | DeepEP | TP=2, EP=64, PP=8, VPP=4 |
| DSV3 | GB200/GB300 | HybridEP | TP=1, EP=64, PP=4, VPP=4 |
| Qwen3 235B | H100 | alltoall + overlap in the current canonical recipe | TP=2, EP=32, PP=8, VPP=4 |
| Qwen3 235B | GB200 | HybridEP | TP=1 or 2, EP=32-64, PP=4, VPP=unspecified |
| Qwen3 30B | 16×H100 | HybridEP | TP=1, EP=16, PP=1, plain EP overlap |
For Qwen3 235B on GB200, explicitly say VPP=unspecified; do not invent or
extrapolate VPP=12 unless a measured row provides it. Treat TE-scoped CUDA
graph scopes (attn, moe_router, moe_preprocess) as profile-driven
candidates,
CUDA_DEVICE_MAX_CONNECTIONS selection,
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, NCCL_GRAPH_REGISTER=0,
GB200/GB300 CPU-side tuning, and the warning not to cargo-cult tracker rows.
These are intentionally rounded so the document stays durable as the tracker moves. Treat them as planning ranges, not exact promises.
| Workload family | Hardware | Typical band | Representative shape |
|---|---|---|---|
| DSV3, large-scale | H100 | low-to-mid hundreds TFLOPS/GPU, high-teens MFU | TP2, EP64, PP8, DeepEP |
| DSV3, large-scale | B200 | high-hundreds TFLOPS/GPU, mid-teens MFU | TP1, EP32, PP8, DeepEP |
| DSV3, large-scale | GB200 | around 1K TFLOPS/GPU, low-20s MFU | TP1, EP64, PP4, HybridEP |
| DSV3, large-scale | GB300 | above the GB200 band, often mid-20s MFU | TP1, EP64, PP4, HybridEP |
| Qwen3 235B | H100 | historical low-300s snapshots; remeasure the current recipe | TP2, EP32, PP8; current recipe uses alltoall + overlap |
| Qwen3 235B | GB200 | high-hundreds TFLOPS/GPU in tuned runs | TP1 or TP2, EP32-64, PP4, HybridEP |
| Qwen3 30B | H100 | about 300 TFLOPS/GPU on the validated 16-GPU shape | TP1, EP16, PP1, HybridEP + EP overlap |
| Qwen3-Next 80B | GB200 | low-300s TFLOPS/GPU in BF16-class runs | TP1, EP32, PP2, HybridEP |
Dispatcher: DeepEP
TP=2 EP=64 PP=8 VPP=4
Routing: force balance
Recompute: light-to-moderate selective recompute
Priority: overlap communication and keep PP efficientDispatcher: DeepEP
TP=1 EP=32 PP=8 VPP=2 or similar
Precision: MXFP8-class
Recompute: selective recompute around MLA up-projection and MLP-side modules
Priority: container quality, PP layout, and DeepEP SMS tuningDispatcher: HybridEP
TP=1 EP=64 PP=4 VPP=4
Precision: MXFP8-class
CUDA Graph: attn + moe_router + moe_preprocess
Priority: HybridEP, CPU optimization, and graph-friendly static shapesDispatcher: alltoall in the current canonical recipe; re-screen flex backends on the target stack
TP=2 EP=32 PP=8 VPP=4
Recompute: none in the current canonical recipe
Priority: communication overlap and router-path cleanupDispatcher: HybridEP
TP=1 or 2 EP=32 to 64 PP=4 VPP=unspecified unless measured
CUDA Graph: attn + moe_router + moe_preprocess
Recompute: moe_act, mlp, or norm depending on memory pressure
Priority: balance throughput against memory headroomDispatcher: HybridEP
TP=1 EP=16 PP=1 CP=1
Precision: BF16
Sequence: 4096
Batch: MBS1 GBS1024
Routing: force balance
EP overlap: enabled
Delayed wgrad: disabled
CUDA Graph: moe_router + moe_preprocess
HybridEP: permute fusion, 32 SMs, 64-token combine chunks
Measured: 20.14729s/step, 299.352 model TFLOPS/GPU over iterations 41-50
Rank-0 peak allocated memory: 62.166 GiBThe current number is the final multi-knob canonical recipe result. An earlier matched A/B isolated plain EP overlap: 244.039 to 287.305 TFLOPS/GPU, with communication hidden by GEMM/attention increasing from 0.11% to 36.55%. Do not attribute the later 299.352 result entirely to overlap.
Dispatcher: HybridEP
TP=1 EP=32 PP=2 VPP around 4
CUDA Graph: attn + moe_router + moe_preprocess
Priority: pipeline layout and grouped GEMM qualityE = embeddingt = transformerm = MTPL = loss| = stage boundaryThe biggest platform difference is usually not just the dispatcher. It is the combination of dispatcher, PP shape, and whether VPP keeps each stage balanced.
| Memory pressure | Starting point |
|---|---|
| low | none or a very narrow selective set |
| moderate | moe_act, mlp, norm, or similar selective modules |
| high | model-specific up-projection plus selective MoE and MLP modules |
| extreme or long-context | full recompute only if the selective path still does not fit |
CUDA_DEVICE_MAX_CONNECTIONS=1
CUDA_DEVICE_MAX_CONNECTIONS=32 # common when EP overlap and CUDA graphs are combined
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
NCCL_GRAPH_REGISTER=0On GB200 and GB300, CPU affinity and general host-overhead cleanup can move the needle almost as much as a dispatcher swap. Treat them as first-class tuning work, not as afterthoughts.
Do not cargo-cult a tracker row: the winning config usually depends on routing mode, container, and PP layout as much as on hardware name.
Container quality matters: large regressions can come from the software stack rather than the model recipe.
VPP must be intentional: a bad VPP split can erase the gain from a better dispatcher.
Compare absolute throughput, not only MFU: MFU can mislead when switching between BF16, FP8, and other precision modes.
Force-balance routing is benchmark-only: it can control routing variance, but it changes semantics. Keep routing fixed within an A/B and validate natural routing separately for training acceptance.
Do not treat the dispatcher table as a hard platform rule: HybridEP is
the validated winner for the canonical 16×H100 Qwen3 30B shape, while the
current 256×H100 Qwen3 235B recipe uses alltoall. Benchmark backend
compatibility and throughput in the production container.
Separate screening, causality, and acceptance: short runs reject weak candidates, matched one-variable A/Bs explain a mechanism, and a 50-step final run validates the complete winner.
Last signature refresh: 2026-08-03.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files in skills/nemo-mbridge-perf-moe-hardware-configs of NVIDIA/skills.
Open the folder on GitHubat commit 14a98ae
Nemo Mbridge Perf Moe Hardware Configs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemo Mbridge Perf Moe Hardware Configs this skillNVIDIA/skills | 3.6k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Add New ModelJakeATX/llamAmpere | 157 | — | ~4.1k | Automated safety check: Pass | MIT | |
| Debug Cuda Crashsgl-project/sglang | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Code ReviewJakeATX/llamAmpere | 157 | — | ~5.6k | Automated safety check: Pass | MIT | |
| Add Modelsohu-mptc/FlashRec | 107 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | |
| AppJakeATX/llamAmpere | 157 | — | ~146 | Automated safety check: Pass | MIT |
JakeATX/llamAmpere
Guided workflow for adding a new model architecture to llama.cpp.
sgl-project/sglang
Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging
JakeATX/llamAmpere
Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR.
sohu-mptc/FlashRec
给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。
JakeATX/llamAmpere
Opinionated app components building on top of ./ui primitives
zhongkaifu/TensorSharp
Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Categories
Representative, point-in-time MoE training playbooks by hardware and model family. Nemo Mbridge Perf Moe Hardware Configs is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Representative, point-in-time MoE training playbooks by hardware and model family.
Nemo Mbridge Perf Moe Hardware Configs fits situations like: AI & LLM Engineering work in your project.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a claude-code`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-hardware-configs in NVIDIA/skills) into .claude/skills/nemo-mbridge-perf-moe-hardware-configs in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a codex`. Or copy the skill folder (skills/nemo-mbridge-perf-moe-hardware-configs in NVIDIA/skills) into .agents/skills/nemo-mbridge-perf-moe-hardware-configs in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemo-mbridge-perf-moe-hardware-configs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemo-mbridge-perf-moe-hardware-configs, .gemini/skills/nemo-mbridge-perf-moe-hardware-configs, .github/skills/nemo-mbridge-perf-moe-hardware-configs and .opencode/skills/nemo-mbridge-perf-moe-hardware-configs in your project.
SKILL.md names no scripts, command-line tools or credentials: Nemo Mbridge Perf Moe Hardware Configs is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Nemo Mbridge Perf Moe Hardware Configs is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Nemo Mbridge Perf Moe Hardware Configs: Add New Model (JakeATX/llamAmpere, 157 stars), Debug Cuda Crash (sgl-project/sglang, 37k stars), Code Review (JakeATX/llamAmpere, 157 stars) and Add Model (sohu-mptc/FlashRec, 107 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,555 GitHub stars. The repository holds 390 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.