Agent skill

Voice Setup

by vellum-ai in vellum-ai/vellum-assistant

Complete voice configuration in chat - desktop Talk and PTT shortcuts, microphone permissions, ElevenLabs/Deepgram TTS, and troubleshooting

MITAuto-check passedMedia & Creative

Install Voice Setup

skills CLI
$ npx skills add vellum-ai/vellum-assistant --skill voice-setup -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vellum-ai/vellum-assistant voice-setup --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vellum-ai/vellum-assistant.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/voice-setup .claude/skills/voice-setup && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice-setup
GitHub stars
1.4k
Token cost
~2.8k tokens
SKILL.md length
1,451 words
Files
2 (incl. assets)
Skills in repo
108
Repo updated
First seen
Licence
MIT

At a glance

Complete voice configuration in chat - desktop Talk and PTT shortcuts, microphone permissions, ElevenLabs/Deepgram TTS, and troubleshooting

  • Works in 4 steps: Microphone Permission → Talk and Push-to-Talk Shortcut → Text-to-Speech Voice (Optional) → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers Available Tools, Setup Flow, Troubleshooting Decision Trees and Deep Debugging, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Voice Setup is an agent skill from vellum-ai/vellum-assistant. Complete voice configuration in chat - desktop Talk and PTT shortcuts, microphone permissions, ElevenLabs/Deepgram TTS, and troubleshooting

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including assets. Compatibility notes: Designed for Vellum personal assistants

It sits in Media & Creative, covering Text to speech and voice. It works with Deepgram, ElevenLabs and macOS. The repository describes itself as: An AI Assistant that’s easy to setup, does your work 24/7, knows your preferences and gets better over time. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “/voice-setup”

Requirements

  • Compatibility (from SKILL.md): Designed for Vellum personal assistants

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Microphone Permission
  2. Talk and Push-to-Talk Shortcut
  3. Text-to-Speech Voice (Optional)
  4. Verification

What it can do on your machine

Read from SKILL.md and the folder at commit 844117a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and powershell).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Vellum personal assistants

    From compatibility in the SKILL.md frontmatter.

Context cost

Voice Setup loads about 2.8k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 1,451 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vellum-ai/vellum-assistant at commit 844117a, republished under its MIT licence (© vellum-ai). 1,451 words, ~2,809 tokens.

Download SKILL.mdSave it as .claude/skills/voice-setup/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
voice-setup
description
Complete voice configuration in chat - desktop Talk and PTT shortcuts, microphone permissions, ElevenLabs/Deepgram TTS, and troubleshooting
compatibility
Designed for Vellum personal assistants
metadata.icon
assets/icon.svg
metadata.emoji
🎙️

You are helping the user set up and troubleshoot voice features entirely within this conversation. Use the client_os: line in <turn_context> to choose the macOS or Windows instructions below. Do not give macOS commands or key names to a Windows user, or Windows guidance to a macOS user.

Before using a desktop tool, check client_os:

  • If it is macos or windows, follow that platform's branch.
  • If it is web, ios, android, or absent, do not call open_system_settings or give desktop shortcut instructions. Explain that permissions and the Talk shortcut must be configured from the Mac or Windows desktop app. You can still complete provider, voice, language, and timeout configuration in the current conversation.

Available Tools

  • voice_config_update changes shared voice settings such as the legacy macOS PTT activation key, conversation timeout, speech providers, and TTS voice ID.
  • open_system_settings opens the correct macOS System Settings or Windows Settings privacy page. Call it only when client_os is macos or windows, and pass that value as platform.
  • navigate_settings_tab opens Vellum settings. Use it for review, or when the desktop-owned Talk shortcut must be recorded in the app.
  • assistant credentials prompt collects API keys securely for ElevenLabs or Deepgram.

The desktop Talk shortcut is client-owned. On current desktop clients, configure it in the Voice settings shortcut control instead of treating voice_config_update setting="activation_key" as a global shortcut editor. Use voice_config_update for the shared settings it owns. Use activation_key only for the legacy macOS PTT activation setting.

Setup Flow

Walk through each relevant section in order. Skip sections the user does not need, and ask before moving to the next section.

1. Microphone Permission

Microphone permission is not reported in <channel_capabilities>. Ask whether Vellum shows a microphone permission warning or whether the user already granted access.

If access is denied or the user is unsure:

  1. Explain that the desktop app needs microphone permission for dictation and voice conversations.
  2. Call open_system_settings with pane: "microphone" and the current platform.
  3. Give the matching instruction:
    • macOS: In System Settings > Privacy & Security > Microphone, turn on Vellum or Vellum Assistant.
    • Windows: In Settings > Privacy & security > Microphone, turn on Microphone access and Let desktop apps access your microphone. Windows groups non-packaged desktop apps under the desktop-app toggle rather than always showing an individual Vellum switch.
  4. Ask the user to return after changing it, then verify with the short recording test in section 4.

If the user confirms access is granted, continue without opening system settings.

2. Talk and Push-to-Talk Shortcut

On macOS, first determine whether the user means the current desktop Talk shortcut or the legacy PTT activation setting. Windows supports the desktop Talk shortcut only.

Desktop Talk shortcut

The Talk shortcut starts or ends a voice conversation.

  • macOS: Offer Fn or a custom global chord. Fn is the Mac-specific helper path and can require Input Monitoring. If macOS refuses Fn registration, direct the user to System Settings > Privacy & Security > Input Monitoring, then have them reopen Vellum. A custom chord should be one the user can spare system-wide.
  • Windows: Fn is unavailable because Windows does not receive the hardware Fn key. Recommend a custom chord for system-wide Talk. The app also offers Ctrl+Shift and Alt taps, but those bare-modifier choices work only while a Vellum window is focused. Warn that another global shortcut can prevent a custom chord from registering.

Ask which behavior they want, then use navigate_settings_tab with tab: "Voice" so they can record the desktop-owned shortcut. Do not claim that voice_config_update changed this shortcut.

Legacy PTT activation setting

This setting is macOS-only. If a Mac user explicitly wants the legacy hold-to-talk setting, offer only values accepted by voice_config_update:

  • fn
  • fn_shift
  • ctrl
  • none

After the user chooses, call voice_config_update with setting: "activation_key" and the matching canonical value.

On Windows, do not offer or call the legacy activation setting. It has no Windows client consumer. Use the desktop Talk shortcut flow instead.

3. Text-to-Speech Voice (Optional)

Ask whether the user wants high-quality text-to-speech voices through ElevenLabs or Deepgram. Standard TTS works without this optional setup.

The included ElevenLabs Voice and Deepgram Voice skills provide the provider-specific setup flow, including API key collection, voice selection, and tuning.

Check the active provider first with assistant config get services.tts.provider. voice_config_update writes the voice to the active provider, and each bring-your-own provider accepts its own voice IDs. If the preferred provider does not match the active provider, collect any required API key before switching:

text
voice_config_update setting="tts_provider" value="deepgram"

The managed vellum provider accepts both supported ElevenLabs and Deepgram voice IDs, so it does not require a provider switch. Then follow the matching included voice skill.

The active provider's voice setting controls both in-app TTS and phone calls.

4. Verification

After setup:

  1. Summarize the configured permissions, shortcut, and provider settings.
  2. Ask the user to test the selected shortcut and speak a short sentence.
  3. Offer to open the Voice settings tab for review with navigate_settings_tab and tab: "Voice".

Desktop Talk starts a live voice session. Its audio is transcribed through the assistant's configured speech-to-text provider over the live voice connection. The Windows native helper provides partials only for one-shot dictation from the microphone button. Ask which surface the user tested before troubleshooting missing text.

Troubleshooting Decision Trees

Show full SKILL.md (591 more words)Show less
"PTT isn't working" or "Can't record"
  1. Ask whether Vellum shows a microphone permission warning and whether the recording indicator appears. If access is denied or capture does not start, follow the microphone permission flow.
  2. Confirm which shortcut the user configured. Legacy PTT applies only on macOS. Verify desktop Talk in the Voice tab.
  3. Apply the platform-specific checks:
    • macOS: Fn requires the native helper and may require Input Monitoring. The Globe key can also be assigned to macOS Dictation or the emoji picker, so suggest a custom chord if both actions fire.
    • Windows: Fn cannot work. A Ctrl+Shift or Alt tap requires Vellum to be focused. For use from another app, record a custom global chord and make sure no other app owns it.
  4. If the user reports Speech Recognition permission as denied or not determined, call open_system_settings with pane: "speech_recognition" and the current platform.
  5. If the Windows shortcut fires but capture does not start, choose Restart from the Vellum tray menu and retry.
"Talk starts but no transcript"
  1. Confirm that the user started Desktop Talk and that its listening or recording indicator appears. If it does not, return to shortcut and microphone capture troubleshooting.
  2. Check the active provider with assistant config get services.stt.provider.
  3. If no usable provider is configured, help the user choose one with voice_config_update setting="stt_provider" and collect any required credential securely before retrying.
  4. If a provider is configured, check for an invalid credential, provider outage, or a live voice connection error. A session that never connects or disconnects before transcript events is a connection path problem, not a Windows recognizer problem.
  5. Confirm the configured speech-to-text language matches the speaker when the selected provider uses that setting. Do not send a Desktop Talk failure to Windows installed-language or native-helper troubleshooting.
"One-shot dictation records but produces no text"

This path applies to the microphone button's one-shot dictation, not Desktop Talk.

  1. If the user reports Speech Recognition permission as denied or unknown, open the platform's privacy page with open_system_settings.
  2. Ask whether dictation partials appear. If capture starts without partials, troubleshoot the native helper and local recognizer.
  3. On Windows, the helper uses an installed Windows speech recognizer and prefers the current Windows display language. If no matching recognizer is installed, add that language's speech feature or select an installed language.
  4. Reduce background noise or move closer to the microphone.
  5. On Windows, choose Restart from the Vellum tray menu and retry. This relaunches Vellum and its one-shot dictation helper.
"Changed a setting but it didn't work"
  1. Shared voice_config_update changes should apply immediately. Verify the persisted value with the relevant config command.
  2. Desktop shortcut changes are stored by the desktop client. Reopen the Voice tab and confirm the Talk shortcut shown there.
  3. If the displayed value is correct but behavior is stale, restart Vellum and retry.

Deep Debugging

For persistent issues, use the matching log path.

macOS:

bash
log stream --predicate 'subsystem == "com.vellum.assistant"' --level debug

Look for voice and speech categories.

Windows PowerShell:

powershell
$log = Get-ChildItem "$env:APPDATA\Vellum*\logs\vellum.log" |
  Sort-Object LastWriteTime -Descending |
  Select-Object -First 1
Get-Content $log.FullName -Wait

For Desktop Talk, look for live voice, WebSocket, and speech-to-text provider errors. For one-shot dictation, look for [win-helper], dictation, permission, and recognizer messages. Do not ask the user to share transcript contents from logs.

Rules

  • Handle every tool-backed setting conversationally in chat.
  • Use navigate_settings_tab for review and for the desktop-owned Talk shortcut, which must be recorded in the client.
  • Use platform-specific Settings names and shortcuts.
  • Be concise. Present the common choices and let the user ask for more.
  • If permission is denied, acknowledge it and explain which voice features will remain unavailable.

© vellum-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (assets) in skills/voice-setup of vellum-ai/vellum-assistant.

  • SKILL.md
  • assets/icon.svg

Open the folder on GitHubat commit 844117a

Compare with similar skills

Voice Setup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice Setup compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice Setup this skillvellum-ai/vellum-assistant1.4k—~2.8kAutomated safety check: PassMIT
9Router Text to Speechdecolua/9router30k—~765Automated safety check: PassMIT
Keirouter Ttsmydisha/keirouter147—~599Automated safety check: PassMIT
Voice AI Developmentdavila7/claude-code-templates32k5 repos~2.1kAutomated safety check: PassMIT
C Voicedaxaur/openpaw174—~401Automated safety check: PassMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • 9Router Text to Speech

    decolua/9router

    Turns text into spoken audio through a 9Router server, choosing among voices from OpenAI, ElevenLabs, Deepgram, Edge TTS and other providers.

    30k GitHub stars~765 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Keirouter Tts

    mydisha/keirouter

    Text-to-speech via KeiRouter /v1/audio/speech using OpenAI / ElevenLabs / Deepgram / Edge TTS / Google TTS / Inworld voices.

    147 GitHub stars~599 tokensUpdated 28 days ago
    Media & CreativeAuto-check passed
  • Voice AI Development

    davila7/claude-code-templates

    Expert in building voice AI applications - from real-time voice agents to voice-enabled apps.

    32k GitHub starsUsed in 5 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • C Voice

    daxaur/openpaw

    Convert speech to text using sag (ElevenLabs STT) and synthesize speech using say (macOS built-in TTS).

    174 GitHub stars~401 tokensUpdated 4 mo ago
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes

More from vellum-ai/vellum-assistant

All 108 skills in this repo
  • Vellum GitHub App Setup

    vellum-ai/vellum-assistant

    Create and configure a GitHub App so the assistant can push commits, open PRs, and comment under its own bot identity.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Discord App Setup

    vellum-ai/vellum-assistant

    Connect a Discord bot to the assistant via the Discord Gateway with guided application creation and intent configuration

    1.4k GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Sentry App Setup

    vellum-ai/vellum-assistant

    Create and configure a Sentry internal integration so the assistant can manage issues, alerts, and releases under its own identity

    1.4k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Memory Corpus Ingest

    vellum-ai/vellum-assistant

    Ingest a large dataset into memory as a skimmed map. An agent skill from vellum-ai/vellum-assistant.

    1.4k GitHub stars~3k tokensUpdated today
    Auto-check: notes
  • Plugin Builder

    vellum-ai/vellum-assistant

    A skill your agent uses when the user wants to build, scaffold, ship, or edit a Vellum plugin that bundles multiple surfaces (hooks, tools, skills, and more) into one installable package.

    1.4k GitHub stars~3.1k tokensUpdated today
    Auto-check passed
  • Slack App Setup

    vellum-ai/vellum-assistant

    Connect a Slack app to the Vellum Assistant via Socket Mode.

    1.4k GitHub stars~2.5k tokensUpdated today
    Auto-check: warnings

Questions about Voice Setup

What does Voice Setup do?

Complete voice configuration in chat - desktop Talk and PTT shortcuts, microphone permissions, ElevenLabs/Deepgram TTS, and troubleshooting. Voice Setup is an agent skill from vellum-ai/vellum-assistant.

When should I use Voice Setup?

Voice Setup fits situations like: tasks that involve Text to speech and voice.

How do I install Voice Setup in Claude Code?

Run `npx skills add vellum-ai/vellum-assistant --skill voice-setup -a claude-code`. Or copy the skill folder (skills/voice-setup in vellum-ai/vellum-assistant) into .claude/skills/voice-setup in your project. Claude Code loads it when a task matches its description.

How do I install Voice Setup in Codex?

Run `npx skills add vellum-ai/vellum-assistant --skill voice-setup -a codex`. Or copy the skill folder (skills/voice-setup in vellum-ai/vellum-assistant) into .agents/skills/voice-setup in your project. Codex loads it when a task matches its description.

Can I use Voice Setup in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vellum-ai/vellum-assistant --skill voice-setup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice-setup, .gemini/skills/voice-setup, .github/skills/voice-setup and .opencode/skills/voice-setup in your project.

What does Voice Setup need to run?

SKILL.md names no scripts, command-line tools or credentials: Voice Setup is instructions for the agent only. Compatibility (from SKILL.md): Designed for Vellum personal assistants.

Does Voice Setup access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Voice Setup safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voice Setup use?

Voice Setup is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voice Setup use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice Setup?

Skills that share tags, products or a category with Voice Setup: 9Router Text to Speech (decolua/9router, 30k stars), Keirouter Tts (mydisha/keirouter, 147 stars), Voice AI Development (davila7/claude-code-templates, 32k stars) and C Voice (daxaur/openpaw, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice Setup?

vellum-ai (a GitHub organization) maintains it in vellum-ai/vellum-assistant, which has 1,400 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 9, 2026.

Source: vellum-ai/vellum-assistant on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.