Nemo
majiayu000/claude-skill-registry
A skill your agent uses for NVIDIA NeMo Speech ASR, TTS, audio processing, SpeechLM2 and voice-agent workflows, speaker diarization, data tooling, and repository development.
Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.
$ npx skills add NVIDIA/skills --skill nemotron-speech -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills nemotron-speech --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemotron-speech .claude/skills/nemotron-speech && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "nemotron-speech" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemotron-speech into .claude/skills/nemotron-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemotron-speech", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/nemotron-speechType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill nemotron-speech -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills nemotron-speech --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/nemotron-speech .agents/skills/nemotron-speech && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "nemotron-speech" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemotron-speech into .agents/skills/nemotron-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemotron-speech", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemotron-speech -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills nemotron-speech --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/nemotron-speech .cursor/skills/nemotron-speech && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "nemotron-speech" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemotron-speech into .cursor/skills/nemotron-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemotron-speech", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/nemotron-speech--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill nemotron-speech -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills nemotron-speech --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/nemotron-speech .gemini/skills/nemotron-speech && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "nemotron-speech" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemotron-speech into .gemini/skills/nemotron-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemotron-speech", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills nemotron-speechInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill nemotron-speech -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/nemotron-speech .github/skills/nemotron-speech && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "nemotron-speech" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemotron-speech into .github/skills/nemotron-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemotron-speech", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill nemotron-speech -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills nemotron-speech --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/nemotron-speech .opencode/skills/nemotron-speech && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "nemotron-speech" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/nemotron-speech into .opencode/skills/nemotron-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "nemotron-speech", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
nemotron-speechRoutes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.
Nemotron Speech is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including scripts and reference files (for example `BENCHMARK.md`, `evals/EVAL.md` and `evals/evals.json`).
It sits in Media & Creative, covering Speech recognition and synthesis and Text to speech and voice. It works with NVIDIA AI Platform and Docker. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pipdockerFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
build.nvidia.comAlso links to:
docs.nvidia.comcatalog.ngc.nvidia.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NVIDIA_API_KEYNGC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Nemotron Speech loads about 3k tokens when it runs, and up to ~47k if it reads all its reference files. Until then it costs about 39 tokens; SKILL.md has 1,049 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
nternally), not just the host user. Run `sudo chown 1000:1000 $LOCAL_NIM_CACHE` after creating the directory so the contAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,049 words, ~2,977 tokens.
.claude/skills/nemotron-speech/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.Note: "Nemotron Speech" is the public-facing name for what NVIDIA documents today as Riva / Riva NIM. All commands, container images, gRPC APIs, Python imports, and documentation URLs still use "Riva" — the rename is brand-only. Do not rename commands, images, or doc URLs.
Agent: When walking the user through a multi-step workflow, announce each step before presenting it: Step N/M — Step Title (e.g., "Step 1/4 — Deploy the Container").
Single entry point for all NVIDIA Nemotron Speech (Riva) NIM workflows: ASR (speech-to-text), TTS (text-to-speech), and NMT (translation). Covers cloud-hosted inference via build.nvidia.com, self-hosted Docker deployment, client-protocol choice for ASR (gRPC, HTTP, WebSocket), custom NeMo model deployment via riva-build, ASR pipeline tuning (VAD, diarization, language models), and the prerequisite Docker / NGC / driver setup.
Use this skill for any Nemotron Speech / Riva NIM task — deployment, testing, custom model build, system requirements check, or model selection across ASR / TTS / NMT modalities.
Identify the user's task type, then load the corresponding reference file from references/. The reference files contain the detailed per-workflow content; this SKILL.md is a routing surface. Load only the reference relevant to the task at hand.
references/setup.md.pip install -U nvidia-riva-client and a valid NVIDIA_API_KEY from https://build.nvidia.com.NVIDIA_API_KEY and NGC_API_KEY as secrets: never print, paste, commit, or log real key values. Prefer --password-stdin for Docker login and store persistent keys in a credential manager or a chmod 600 env file rather than world-readable shell startup files./opt/nim/.cache must be writable by the container user (the NIM container runs as nvs:1000 internally), not just the host user. Run sudo chown 1000:1000 $LOCAL_NIM_CACHE after creating the directory so the container can write to it. Avoid world-writable modes — they let any local user replace cached model artifacts. Also avoid -u "$(id -u):$(id -g)" on the docker run — /opt/nim/workspace inside the container isn't writable to arbitrary UIDs. If you see I/O error Permission denied (os error 13) during model download, the host directory ownership is the issue.references/setup.md.references/deployment-readiness-checks.md.references/model-selection.md.references/asr.md..nemo → RMIR → NIM) to references/asr-custom.md.references/pipelines.md.references/tts.md..nemo → RMIR → NIM) to references/tts-custom.md.custom_configuration keys — to references/tts-pipelines.md. (For discovering and constructing the pronunciation itself, use tts-pronunciation.md.)references/tts-pronunciation.md.references/nmt.md.For per-release detail — current model catalog, container IDs, function IDs, voice lists, VRAM minimums, per-model feature support — fetch or open the canonical NVIDIA doc rather than relying on text in this SKILL.md or the references. Each reference file includes its own routing table to the relevant doc pages.
Top-level landing pages:
"Deploy a Parakeet ASR NIM" → load references/asr.md, follow Option B (self-hosted), Steps 1–4.
"Synthesize speech with Magpie" → load references/tts.md, follow Option A (cloud) or Option B (self-hosted).
"Translate English to German" → load references/nmt.md, follow the 4-step flow.
"Convert my fine-tuned .nemo to a NIM" → load references/asr-custom.md for the 4-phase pipeline and references/pipelines.md for build-time config.
"Deploy a custom fine-tuned TTS voice as a NIM" → load references/tts-custom.md for the 4-phase pipeline.
"Use zero-shot voice cloning with Magpie" → load references/tts-pipelines.md.
"Add SSML emphasis tags to my TTS request" → load references/tts-pipelines.md.
"'NVIDIA' sounds wrong in my Magpie TTS output — suggest a few IPA options to test" → load references/tts-pronunciation.md, generate IPA candidates, synthesize variants, then output all three delivery formats.
"How do I fix the pronunciation of 'NIM' in Riva TTS with a custom_dictionary in gRPC Python?" → load references/tts-pronunciation.md, propose IPA for 'NIM', show wire format and gRPC snippet.
"Can my GPU run this?" → load references/deployment-readiness-checks.md and run the 6-step system check.
"Which Riva model should I use?" → load references/model-selection.md, apply the decision framework, then fetch the support matrix for the specific current model name.
riva-build, riva-deploy, riva_streaming_asr_client), Python client (riva.client), gRPC namespace (nvidia.riva.asr.*), container registry (nvcr.io/nim/nvidia/*), and all NVIDIA documentation URLs still use "Riva". Do not rename these in code, commands, or docs.For task-specific runtime or modality issues, use the relevant reference file (references/<task>.md). Cross-cutting readiness checks:
references/deployment-readiness-checks.md (system check + health check table)references/deployment-readiness-checks.mddocker pull from nvcr.io returns 403 → references/setup.md (Step 5 — Docker login)references/asr-custom.md (Phase 2 base image)references/deployment-readiness-checks.md, then verify on the support matrixreferences/setup.md)NVIDIA_API_KEY and internet accessriva.client), gRPC services (nvidia.riva.*), and NVIDIA documentation URLs still use "Riva" — follow official docs and catalogs for naming, do not rename these in commands or codereferences/deployment-readiness-checks.mdreferences/setup.mdreferences/model-selection.mdreferences/asr.md, references/tts.md, or references/nmt.md© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 17 other files (scripts, references) in skills/nemotron-speech of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Nemotron Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Nemotron Speech this skillNVIDIA/skills | 3.5k | — | ~3k | Automated safety check: Notes | Apache-2.0 | |
| Nemomajiayu000/claude-skill-registry | 666 | 1 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| 9Router Speech-to-Textdecolua/9router | 30k | — | ~745 | Automated safety check: Pass | MIT | |
| Dotty Av TestBrettKinny/dotty-stackchan | 113 | — | ~866 | Automated safety check: Notes | MIT | |
| Local AI Useamd/skills | 398 | — | ~5k | Automated safety check: Notes | MIT | |
| Arkcli Code Examplevolcengine/ark-cli | 140 | — | ~743 | Automated safety check: Pass | Apache-2.0 |
majiayu000/claude-skill-registry
A skill your agent uses for NVIDIA NeMo Speech ASR, TTS, audio processing, SpeechLM2 and voice-agent workflows, speaker diarization, data tooling, and repository development.
decolua/9router
Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.
BrettKinny/dotty-stackchan
Run local black-box voice tests against the physical Dotty robot by playing a TTS prompt through the workstation speakers while the C920 records video and room audio.
amd/skills
Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.
volcengine/ark-cli
arkcli +code-example:为指定基础模型生成多语言(Python / Go / Java / Node / curl)调用示例代码并写入本地文件。数据源是火山方舟 OpenTOP OpenGetSampleCode。当用户需要拿某个基础模型的 SDK / curl 调用示例、保存为本地接入模板时使用。反触发:TTS/ASR/语音模型没有 arkcli…
davila7/claude-code-templates
Expert in building voice AI applications - from real-time voice agents to voice-enabled apps.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted. Nemotron Speech is an agent skill from NVIDIA/skills, published by the product's own GitHub organization.com or self-hosted.
Nemotron Speech fits situations like: tasks that involve Speech recognition and synthesis; tasks that involve Text to speech and voice.
Run `npx skills add NVIDIA/skills --skill nemotron-speech -a claude-code`. Or copy the skill folder (skills/nemotron-speech in NVIDIA/skills) into .claude/skills/nemotron-speech in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill nemotron-speech -a codex`. Or copy the skill folder (skills/nemotron-speech in NVIDIA/skills) into .agents/skills/nemotron-speech in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemotron-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemotron-speech, .gemini/skills/nemotron-speech, .github/skills/nemotron-speech and .opencode/skills/nemotron-speech in your project.
Going by SKILL.md and its folder, Nemotron Speech needs Python for the scripts in its folder, the command-line tools its instructions call (pip and docker) and credentials named NVIDIA_API_KEY and NGC_API_KEY. Our summary lists: Python 3; Docker; A credential in NVIDIA_API_KEY; A credential in NGC_API_KEY.
SKILL.md names 3 domains. In commands or code: build.nvidia.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.nvidia.com and catalog.ngc.nvidia.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Nemotron Speech is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 44k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Nemotron Speech: Nemo (majiayu000/claude-skill-registry, 666 stars), 9Router Speech-to-Text (decolua/9router, 30k stars), Dotty Av Test (BrettKinny/dotty-stackchan, 113 stars) and Local AI Use (amd/skills, 398 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.