Official agent skill

Nemotron Voice Agent Builder

by NVIDIA in NVIDIA/skills

Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Nemotron Voice Agent Builder

skills CLI
$ npx skills add NVIDIA/skills --skill nemotron-voice-agent-builder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills nemotron-voice-agent-builder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/nemotron-voice-agent-builder .claude/skills/nemotron-voice-agent-builder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
nemotron-voice-agent-builder
GitHub stars
3.5k
Token cost
~844 tokens
SKILL.md length
356 words
Files
32 (incl. references)
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit.

  • Iterating on a real-time voice-agent pipeline
  • SKILL.md covers When to Use This Skill, Workflow and Rules
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Including speech (ASR/TTS) customization and cloud

What it does

Nemotron Voice Agent Builder is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. Use when building, scaffolding, or iterating on a real-time voice-agent pipeline, including speech (ASR/TTS) customization and cloud or local deployment. Not for offline/batch speech-to-text, text-only chat or RAG, or generic Docker or CUDA work unrelated to a voice agent.

Its SKILL.md is about 840 tokens, which your agent loads only when the skill is triggered. The skill folder holds 36 other files, including reference files (for example `BENCHMARK.md`, `evals/EVAL.md` and `evals/evals.json`).

It sits in AI & LLM Engineering, covering Speech recognition and synthesis. It works with NVIDIA AI Platform, CUDA and Docker. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Iterating on a real-time voice-agent pipeline
  • Including speech (ASR/TTS) customization and cloud
  • Local deployment

Example prompts

  • “/nemotron-voice-agent-builder”

Requirements

  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Nemotron Voice Agent Builder loads about 844 tokens when it runs, and up to ~47k if it reads all its reference files. Until then it costs about 97 tokens; SKILL.md has 356 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~844
With references · SKILL.md plus every file in references/, read only if the agent opens them
~47k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 356 words, ~844 tokens.

Download SKILL.mdSave it as .claude/skills/nemotron-voice-agent-builder/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.
name
nemotron-voice-agent-builder
description
Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. Use when building, scaffolding, or iterating on a real-time voice-agent pipeline, including speech (ASR/TTS) customization and cloud or local deployment. Not for offline/batch speech-to-text, text-only chat or RAG, or generic Docker or CUDA work unrelated to a voice agent.
version
2.2.0
license
CC-BY-4.0 AND Apache-2.0
metadata.author
NVIDIA Voice Agent Team <nemotron-voice-agent@nvidia.com>
metadata.tags
voice-agent, nvidia, nemotron, magpie, pipecat, livekit, nim, vllm, nemo-speech, omni

Create Voice Agent

Creates a working NVIDIA voice agent or updates an existing project:

  • Cascaded: ASR transcribes, a text LLM answers, TTS speaks.
  • Omni: one audio-in LLM replaces ASR and the text LLM. TTS still speaks.

When to Use This Skill

Use this skill when the user wants to build, scaffold, configure, refine, or fix an NVIDIA voice agent — a real-time speech pipeline with audio input and spoken output — on Pipecat or LiveKit. This covers Cascaded (ASR → LLM → TTS) and Omni (audio-in LLM + TTS) pipelines, speech customization (ASR word boosting, TTS pronunciation), multilingual routing, and cloud or local deployment (NIM, vLLM, NeMo-Speech.cpp) on workstations, DGX Spark, or Jetson Thor.

Do not use this skill for text-only chatbots or RAG, standalone batch speech-to-text transcription, generic Docker or infrastructure help, or unrelated CUDA or model work that has no voice-agent pipeline.

Workflow

Classify the starting state before following the phases:

  • For an empty project, follow all phases in order.
  • For a working existing project, read references/operations/iterate.md first, and then route only the references needed for the requested change.
  • For a broken existing project, read references/operations/troubleshoot.md first. After restoring the baseline, continue with references/operations/iterate.md.

For an empty project, follow these phases in order:

PhaseRead
1. Intakereferences/intake.md resolves pipeline and framework first, then routes preflight, models, language, and domain files
2. Buildreferences/platforms/readiness.md if self-hosted → exact profile or quantization variant in the routed model file → references/output-contract.md → routed deployment path → references/operations/observability.md → selected framework files
3. Hand overreferences/operations/run.md gates the client on the generated scripts/smoke.sh, then references/operations/troubleshoot.md if the spoken exchange fails
4. Iteratereferences/operations/iterate.md for changes to a working agent
Show full SKILL.md (88 more words)Show less

Wait for approval before writing files.

references/preflight.md §4 probes the host and selects the deployment path. The routed platform file owns that path from there.

Also read references/networking/remote-webrtc.md when Pipecat WebRTC crosses hosts or networks.

Rules

  • Resolve every model id, image tag, profile, and serve flag from current NVIDIA documentation at build time. Never from memory.
  • Read every routed file before generating code, and keep the language and behavior locks approved in intake.
  • Handover requires a successful spoken exchange. A running process is not a working voice agent.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 31 other files (references) in skills/nemotron-voice-agent-builder of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/EVAL.md
  • evals/evals.json
  • references/domain/agent-behavior.md
  • references/domain/speech-customization.md
  • references/frameworks/livekit.md
  • references/frameworks/omni.md
  • references/frameworks/pipecat.md
  • references/intake.md
  • references/models/asr.md
  • references/models/catalog.md
  • references/models/language-routing.md
  • references/models/llm.md
  • references/models/tts.md
  • references/networking
  • … and 16 more

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Nemotron Voice Agent Builder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Nemotron Voice Agent Builder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Nemotron Voice Agent Builder this skillNVIDIA/skills3.5k—~844Automated safety check: PassApache-2.0
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.0
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Init GPU Serverdrawthingsai/draw-things-community582—~2.2kAutomated safety check: PassGPL-3.0
Dstack Prototypingdstackai/dstack2.3k—~1.6kAutomated safety check: PassMPL-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0

Similar skills

  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    146 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Init GPU Server

    drawthingsai/draw-things-community

    Initialize a Draw Things GPU server with GPUScript, including script sync, Docker/CUDA/NVIDIA runtime setup, 7T data disk mounting, mergerfs, and end-to-end GPU verification.

    582 GitHub stars~2.2k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Dstack Prototyping

    dstackai/dstack

    Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven.

    2.3k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Nemotron Voice Agent Builder

What does Nemotron Voice Agent Builder do?

Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit. Nemotron Voice Agent Builder is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Create, refine, or fix NVIDIA voice agents (Cascaded or Omni) with Pipecat or LiveKit.

When should I use Nemotron Voice Agent Builder?

Nemotron Voice Agent Builder fits situations like: iterating on a real-time voice-agent pipeline; including speech (ASR/TTS) customization and cloud; local deployment.

How do I install Nemotron Voice Agent Builder in Claude Code?

Run `npx skills add NVIDIA/skills --skill nemotron-voice-agent-builder -a claude-code`. Or copy the skill folder (skills/nemotron-voice-agent-builder in NVIDIA/skills) into .claude/skills/nemotron-voice-agent-builder in your project. Claude Code loads it when a task matches its description.

How do I install Nemotron Voice Agent Builder in Codex?

Run `npx skills add NVIDIA/skills --skill nemotron-voice-agent-builder -a codex`. Or copy the skill folder (skills/nemotron-voice-agent-builder in NVIDIA/skills) into .agents/skills/nemotron-voice-agent-builder in your project. Codex loads it when a task matches its description.

Can I use Nemotron Voice Agent Builder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill nemotron-voice-agent-builder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/nemotron-voice-agent-builder, .gemini/skills/nemotron-voice-agent-builder, .github/skills/nemotron-voice-agent-builder and .opencode/skills/nemotron-voice-agent-builder in your project.

What does Nemotron Voice Agent Builder need to run?

SKILL.md names no scripts, command-line tools or credentials: Nemotron Voice Agent Builder is instructions for the agent only. Our summary lists: Docker.

Does Nemotron Voice Agent Builder access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Nemotron Voice Agent Builder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Nemotron Voice Agent Builder use?

Nemotron Voice Agent Builder is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Nemotron Voice Agent Builder use?

About 844 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 46k tokens, read only when the agent opens those files.

What are the alternatives to Nemotron Voice Agent Builder?

Skills that share tags, products or a category with Nemotron Voice Agent Builder: Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 146 stars), Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars), Init GPU Server (drawthingsai/draw-things-community, 582 stars) and Dstack Prototyping (dstackai/dstack, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Nemotron Voice Agent Builder?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.