Official agent skill

Nemotron Speech

by NVIDIA in NVIDIA/skills

Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.

OfficialApache-2.0Auto-check: notesMedia & Creative

Install Nemotron Speech

skills CLI
$ npx skills add NVIDIA/skills --skill nemotron-speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemotron-speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemotron-speech .claude/skills/nemotron-speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemotron-speech
GitHub stars
3.5k
Token cost
~3k tokens
SKILL.md length
1,049 words
Files
18 (incl. scripts, references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.

  • Tasks that involve Speech recognition and synthesis
  • SKILL.md covers Purpose, When to Use This Skill, Workflow and Prerequisites, plus 7 more sections
  • Runs Python scripts from its folder; calls pip and docker; reaches build.nvidia.com; needs NVIDIA_API_KEY and NGC_API_KEY
  • Tasks that involve Text to speech and voice

What it does

Nemotron Speech is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including scripts and reference files (for example `BENCHMARK.md`, `evals/EVAL.md` and `evals/evals.json`).

It sits in Media & Creative, covering Speech recognition and synthesis and Text to speech and voice. It works with NVIDIA AI Platform and Docker. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Speech recognition and synthesis
  • Tasks that involve Text to speech and voice

Example prompts

  • “Use the nemotron-speech skill to route NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on…”
  • “/nemotron-speech”

Requirements

  • Python 3
  • Docker
  • A credential in NVIDIA_API_KEY
  • A credential in NGC_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • build.nvidia.com

    Also links to:

    • docs.nvidia.com
    • catalog.ngc.nvidia.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NVIDIA_API_KEY
    • NGC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemotron Speech loads about 3k tokens when it runs, and up to ~47k if it reads all its reference files. Until then it costs about 39 tokens; SKILL.md has 1,049 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~47k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:102
    nternally), not just the host user. Run `sudo chown 1000:1000 $LOCAL_NIM_CACHE` after creating the directory so the cont

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,049 words, ~2,977 tokens.

Download SKILL.mdSave it as .claude/skills/nemotron-speech/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
nemotron-speech
description
Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted.
triggers
Nemotron Speech, deploy Riva NIM, deploy ASR/TTS/NMT NIM, Riva ASR, Riva TTS, Riva translation, Parakeet, Canary, Whisper, Nemotron ASR Streaming, Magpie TTS…
version
1.0.0
license
Apache-2.0
metadata.author
Nemotron Speech Team
metadata.team
riva
metadata.tags
nvidia, nemotron-speech, riva, nim, asr, tts, nmt, speech, speech-to-text, text-to-speech, translation, parakeet, canary, whisper, magpie, nemotron, grpc…
metadata.domain
ml

Nemotron Speech Skills

Note: "Nemotron Speech" is the public-facing name for what NVIDIA documents today as Riva / Riva NIM. All commands, container images, gRPC APIs, Python imports, and documentation URLs still use "Riva" — the rename is brand-only. Do not rename commands, images, or doc URLs.

Agent: When walking the user through a multi-step workflow, announce each step before presenting it: Step N/M — Step Title (e.g., "Step 1/4 — Deploy the Container").

Purpose

Single entry point for all NVIDIA Nemotron Speech (Riva) NIM workflows: ASR (speech-to-text), TTS (text-to-speech), and NMT (translation). Covers cloud-hosted inference via build.nvidia.com, self-hosted Docker deployment, client-protocol choice for ASR (gRPC, HTTP, WebSocket), custom NeMo model deployment via riva-build, ASR pipeline tuning (VAD, diarization, language models), and the prerequisite Docker / NGC / driver setup.

When to Use This Skill

Use this skill for any Nemotron Speech / Riva NIM task — deployment, testing, custom model build, system requirements check, or model selection across ASR / TTS / NMT modalities.

Workflow

Identify the user's task type, then load the corresponding reference file from references/. The reference files contain the detailed per-workflow content; this SKILL.md is a routing surface. Load only the reference relevant to the task at hand.

Prerequisites

  • For self-hosted deployment: NVIDIA AI Enterprise (NVAIE) entitlement, then complete the environment setup — NVIDIA drivers, Docker, Container Toolkit, NGC API key, Riva Python client. See references/setup.md.
  • For cloud-hosted inference: pip install -U nvidia-riva-client and a valid NVIDIA_API_KEY from https://build.nvidia.com.
  • Treat NVIDIA_API_KEY and NGC_API_KEY as secrets: never print, paste, commit, or log real key values. Prefer --password-stdin for Docker login and store persistent keys in a credential manager or a chmod 600 env file rather than world-readable shell startup files.
  • For self-hosted Docker model caching: host directories mounted at /opt/nim/.cache must be writable by the container user (the NIM container runs as nvs:1000 internally), not just the host user. Run sudo chown 1000:1000 $LOCAL_NIM_CACHE after creating the directory so the container can write to it. Avoid world-writable modes — they let any local user replace cached model artifacts. Also avoid -u "$(id -u):$(id -g)" on the docker run — /opt/nim/workspace inside the container isn't writable to arbitrary UIDs. If you see I/O error Permission denied (os error 13) during model download, the host directory ownership is the issue.

Instructions

  • Match the user's task to one reference file and load only that file; the references are detailed, so progressive disclosure keeps context tight.
  • Route setup requests for drivers, Docker, Container Toolkit, and NGC to references/setup.md.
  • Route GPU compatibility, deployment readiness, and container health checks to references/deployment-readiness-checks.md.
  • Route model choice across ASR, TTS, and NMT to references/model-selection.md.
  • Route ASR deployment or inference for Parakeet, Canary, Whisper, and Nemotron ASR Streaming to references/asr.md.
  • Route custom-trained NeMo ASR deployment (.nemo → RMIR → NIM) to references/asr-custom.md.
  • Route ASR pipeline configuration for VAD, diarization, language models, and chunk size to references/pipelines.md.
  • Route TTS deployment or inference for Magpie to references/tts.md.
  • Route custom or fine-tuned TTS model deployment (.nemo → RMIR → NIM) to references/tts-custom.md.
  • Route TTS synthesis pipeline configuration — SSML, zero-shot voice cloning, applying an existing pronunciation dictionary, audio encoding, sample rate, custom_configuration keys — to references/tts-pipelines.md. (For discovering and constructing the pronunciation itself, use tts-pronunciation.md.)
  • Route TTS pronunciation discovery — finding, testing, and applying IPA pronunciations for specific words or phrases — to references/tts-pronunciation.md.
  • Route NMT deployment or inference for Riva Translate, language pairs, and DNT tags to references/nmt.md.

Source of truth

For per-release detail — current model catalog, container IDs, function IDs, voice lists, VRAM minimums, per-model feature support — fetch or open the canonical NVIDIA doc rather than relying on text in this SKILL.md or the references. Each reference file includes its own routing table to the relevant doc pages.

Top-level landing pages:

Show full SKILL.md (377 more words)Show less

Examples

"Deploy a Parakeet ASR NIM" → load references/asr.md, follow Option B (self-hosted), Steps 1–4.

"Synthesize speech with Magpie" → load references/tts.md, follow Option A (cloud) or Option B (self-hosted).

"Translate English to German" → load references/nmt.md, follow the 4-step flow.

"Convert my fine-tuned .nemo to a NIM" → load references/asr-custom.md for the 4-phase pipeline and references/pipelines.md for build-time config.

"Deploy a custom fine-tuned TTS voice as a NIM" → load references/tts-custom.md for the 4-phase pipeline.

"Use zero-shot voice cloning with Magpie" → load references/tts-pipelines.md.

"Add SSML emphasis tags to my TTS request" → load references/tts-pipelines.md.

"'NVIDIA' sounds wrong in my Magpie TTS output — suggest a few IPA options to test" → load references/tts-pronunciation.md, generate IPA candidates, synthesize variants, then output all three delivery formats.

"How do I fix the pronunciation of 'NIM' in Riva TTS with a custom_dictionary in gRPC Python?" → load references/tts-pronunciation.md, propose IPA for 'NIM', show wire format and gRPC snippet.

"Can my GPU run this?" → load references/deployment-readiness-checks.md and run the 6-step system check.

"Which Riva model should I use?" → load references/model-selection.md, apply the decision framework, then fetch the support matrix for the specific current model name.

Naming & Terminology

  • Skill brand: Nemotron Speech (public-facing name).
  • Internal naming preserved: commands (riva-build, riva-deploy, riva_streaming_asr_client), Python client (riva.client), gRPC namespace (nvidia.riva.asr.*), container registry (nvcr.io/nim/nvidia/*), and all NVIDIA documentation URLs still use "Riva". Do not rename these in code, commands, or docs.

Troubleshooting

For task-specific runtime or modality issues, use the relevant reference file (references/<task>.md). Cross-cutting readiness checks:

Limitations

  • x86_64 architecture only — WSL2 on Windows requires Podman and supports a subset of NIMs (see references/setup.md)
  • Self-hosted deployment requires an NVIDIA AI Enterprise license
  • Cloud-hosted inference requires an active NVIDIA_API_KEY and internet access
  • Public skill branding is "Nemotron Speech"; commands, container images, Python imports (riva.client), gRPC services (nvidia.riva.*), and NVIDIA documentation URLs still use "Riva" — follow official docs and catalogs for naming, do not rename these in commands or code

Next Steps

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts, references) in skills/nemotron-speech of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/EVAL.md
  • evals/evals.json
  • references/asr-custom.md
  • references/asr.md
  • references/deployment-readiness-checks.md
  • references/model-selection.md
  • references/nmt.md
  • references/pipelines.md
  • references/setup.md
  • references/tts-custom.md
  • references/tts-pipelines.md
  • references/tts-pronunciation.md
  • references/tts.md
  • scripts/main.py
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Nemotron Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemotron Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemotron Speech this skillNVIDIA/skills3.5k—~3kAutomated safety check: NotesApache-2.0
Nemomajiayu000/claude-skill-registry6661 repos~1.5kAutomated safety check: PassApache-2.0
9Router Speech-to-Textdecolua/9router30k—~745Automated safety check: PassMIT
Dotty Av TestBrettKinny/dotty-stackchan113—~866Automated safety check: NotesMIT
Local AI Useamd/skills398—~5kAutomated safety check: NotesMIT
Arkcli Code Examplevolcengine/ark-cli140—~743Automated safety check: PassApache-2.0

Similar skills

  • Nemo

    majiayu000/claude-skill-registry

    A skill your agent uses for NVIDIA NeMo Speech ASR, TTS, audio processing, SpeechLM2 and voice-agent workflows, speaker diarization, data tooling, and repository development.

    666 GitHub starsUsed in 1 repo~1.5k tokens
    AI & LLM EngineeringAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    30k GitHub stars~745 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Dotty Av Test

    BrettKinny/dotty-stackchan

    Run local black-box voice tests against the physical Dotty robot by playing a TTS prompt through the workstation speakers while the C920 records video and room audio.

    113 GitHub stars~866 tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Arkcli Code Example

    volcengine/ark-cli

    arkcli +code-example:为指定基础模型生成多语言(Python / Go / Java / Node / curl)调用示例代码并写入本地文件。数据源是火山方舟 OpenTOP OpenGetSampleCode。当用户需要拿某个基础模型的 SDK / curl 调用示例、保存为本地接入模板时使用。反触发:TTS/ASR/语音模型没有 arkcli…

    140 GitHub stars~743 tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Voice AI Development

    davila7/claude-code-templates

    Expert in building voice AI applications - from real-time voice agents to voice-enabled apps.

    32k GitHub starsUsed in 6 repos~2.1k tokens
    Media & CreativeAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemotron Speech

What does Nemotron Speech do?

Routes NVIDIA Nemotron Speech (Formerly Riva) NIM tasks — deploys, runs, and tests ASR, TTS, and NMT NIMs on build.nvidia.com or self-hosted. Nemotron Speech is an agent skill from NVIDIA/skills, published by the product's own GitHub organization.com or self-hosted.

When should I use Nemotron Speech?

Nemotron Speech fits situations like: tasks that involve Speech recognition and synthesis; tasks that involve Text to speech and voice.

How do I install Nemotron Speech in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemotron-speech -a claude-code`. Or copy the skill folder (skills/nemotron-speech in NVIDIA/skills) into .claude/skills/nemotron-speech in your project. Claude Code loads it when a task matches its description.

How do I install Nemotron Speech in Codex?

Run `npx skills add NVIDIA/skills --skill nemotron-speech -a codex`. Or copy the skill folder (skills/nemotron-speech in NVIDIA/skills) into .agents/skills/nemotron-speech in your project. Codex loads it when a task matches its description.

Can I use Nemotron Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemotron-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemotron-speech, .gemini/skills/nemotron-speech, .github/skills/nemotron-speech and .opencode/skills/nemotron-speech in your project.

What does Nemotron Speech need to run?

Going by SKILL.md and its folder, Nemotron Speech needs Python for the scripts in its folder, the command-line tools its instructions call (pip and docker) and credentials named NVIDIA_API_KEY and NGC_API_KEY. Our summary lists: Python 3; Docker; A credential in NVIDIA_API_KEY; A credential in NGC_API_KEY.

Does Nemotron Speech access the network?

SKILL.md names 3 domains. In commands or code: build.nvidia.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.nvidia.com and catalog.ngc.nvidia.com. This is read from the text; nothing was executed.

Is Nemotron Speech safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Nemotron Speech use?

Nemotron Speech is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemotron Speech use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 44k tokens, read only when the agent opens those files.

What are the alternatives to Nemotron Speech?

Skills that share tags, products or a category with Nemotron Speech: Nemo (majiayu000/claude-skill-registry, 666 stars), 9Router Speech-to-Text (decolua/9router, 30k stars), Dotty Av Test (BrettKinny/dotty-stackchan, 113 stars) and Local AI Use (amd/skills, 398 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemotron Speech?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.