Official agent skill

Grok Realtime Voice Integration

by cursor in cursor/plugins

Wires Grok speech-to-speech into an app's own microphone and audio playback over a realtime WebSocket, replacing an STT-LLM-TTS cascade or OpenAI Realtime.

OfficialNo licenceAuto-check passedMedia & Creative

Install Grok Realtime Voice Integration

skills CLI
$ npx skills add cursor/plugins --skill add-voice -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cursor/plugins add-voice --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cursor/plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/grok-voice/skills/add-voice .claude/skills/add-voice && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-voice
GitHub stars
10k
Token cost
~1.7k tokens
SKILL.md length
643 words
Files
1
Skills in repo
99
Repo updated
First seen
Licence
None found

At a glance

Wires Grok speech-to-speech into an app's own microphone and audio playback over a realtime WebSocket, replacing an STT-LLM-TTS cascade or OpenAI Realtime.

  • Works in 9 steps: Map the app → Auth → Connect + session → …
  • Adding Grok speech-to-speech voice mode to an existing app
  • SKILL.md covers Goal, Protocol first, Docs and Steps, plus 1 more section
  • Reaches api.x.ai; needs XAI_API_KEY

What it does

This skill wires a duplex voice path into an existing app: the app's microphone goes in, audio comes back out, over `wss://api.x.ai/v1/realtime?model=grok-voice-latest`, with safe authentication. Cursor itself has no microphone, so the skill wires the app or a sample client rather than the IDE. It first maps the app's current stack, whether none, an OpenAI Realtime integration, an STT-to-LLM-to-TTS cascade, or standalone TTS or STT, and replaces an existing cascade or OpenAI Realtime with the single duplex loop while keeping standalone one-shot listen or speak endpoints only if the product still needs them outside the agent.

Authentication uses a server-side Bearer `XAI_API_KEY` for a server client, or for browser and mobile clients a backend-issued ephemeral token from a `client_secrets` endpoint, since a long-lived key must never ship in a client bundle. On connect, a `session.update` message sets the voice, instructions and turn detection mode, and `audio.input.transcription.model` must be set to `grok-transcribe` or no user transcript arrives. On the app side, one AudioContext per session captures and plays back 24 kHz PCM audio, created inside a user gesture to satisfy autoplay policy, with microphone audio chunked through an AudioWorklet in roughly 100 millisecond pieces.

When your agent uses it

  • Adding Grok speech-to-speech voice mode to an existing app
  • Replacing an STT-LLM-TTS cascade with a single realtime voice loop
  • Replacing an existing OpenAI Realtime voice integration with Grok

Example prompts

  • “Add Grok voice mode to our app, replacing the current STT-LLM-TTS cascade.”
  • “Wire up Grok realtime speech-to-speech for our web client with ephemeral token auth.”
  • “Switch our OpenAI Realtime voice integration over to Grok voice.”

Requirements

  • An XAI_API_KEY for server-side auth
  • A backend endpoint to mint ephemeral tokens for browser or mobile clients

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Map the app
  2. Auth
  3. Connect + session
  4. Audio I/O (app-side)
  5. Composer UI convention
  6. TS skeleton (default)
  7. Python twin (only if the app is Python)
  8. Instrument (before the first human test)
  9. Smoke

What it can do on your machine

Read from SKILL.md and the folder at commit 9f451cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.x.ai

    Also links to:

    • docs.x.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • XAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grok Realtime Voice Integration loads about 1.7k tokens when it runs. Until then it costs about 105 tokens; SKILL.md has 643 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 643 words (~1,724 tokens).

“Add Grok Speech to Speech to an existing app. Run on /add-voice, typed Voice Mode, or clear “add Grok voice” intent.”

— opening of SKILL.md by cursor
name
add-voice

Read the full SKILL.md on GitHub

Files

Just SKILL.md in grok-voice/skills/add-voice of cursor/plugins.

Open the folder on GitHubat commit 9f451cf

Compare with similar skills

Grok Realtime Voice Integration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grok Realtime Voice Integration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grok Realtime Voice Integration this skillcursor/plugins10k—~1.7kAutomated safety check: PassNone
Deepgram Python Text-to-Speechdeepgram/deepgram-python-sdk469—~1.8kAutomated safety check: PassMIT
Speech Engineelevenlabs/skills479—~2.5kAutomated safety check: WarnMIT
Local AI Useamd/skills395—~5kAutomated safety check: NotesMIT
Voice AI Developmentdavila7/claude-code-templates32k5 repos~2.1kAutomated safety check: PassMIT
Agentselevenlabs/skills479—~6.5kAutomated safety check: PassMIT

Similar skills

  • Deepgram Python Text-to-Speech

    deepgram/deepgram-python-sdk

    Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper.

    469 GitHub stars~1.8k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    479 GitHub stars~2.5k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    395 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Voice AI Development

    davila7/claude-code-templates

    Expert in building voice AI applications - from real-time voice agents to voice-enabled apps.

    32k GitHub starsUsed in 5 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • Agents

    elevenlabs/skills

    Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills.

    479 GitHub stars~6.5k tokensUpdated today
    Backend & APIsAuto-check passed
  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    395 GitHub stars~6k tokensUpdated today
    Media & CreativeAuto-check passed

More from cursor/plugins

All 99 skills in this repo
  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 9 repos~1.6k tokens
    Auto-check passed
  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    Auto-check passed
  • Official

    Starts three parallel reviewer subagents over the current conversation transcript, then turns their findings into concrete edits to existing skills.

    10k GitHub starsUsed in 5 repos~1.2k tokens
    Auto-check passed
  • Official

    Applies four layers of technical-writing rules to docs, RFCs, readmes, PR descriptions and commit messages so a tired engineer follows them on the first read.

    10k GitHub starsUsed in 10 repos~2.4k tokens
    Auto-check passed
  • Advisor Mode

    cursor/plugins

    Official

    Adds a second, stronger model that the main agent consults before major decisions, when stuck and before finishing, controlled by /advisor commands.

    10k GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Official

    Prepare PRs for review by cleaning noisy history, improving PR descriptions, and adding reviewer guidance without changing code behavior.

    10k GitHub starsUsed in 3 repos~569 tokens
    Auto-check passed

Works with

Questions about Grok Realtime Voice Integration

What does Grok Realtime Voice Integration do?

Wires Grok speech-to-speech into an app's own microphone and audio playback over a realtime WebSocket, replacing an STT-LLM-TTS cascade or OpenAI Realtime. model=grok-voice-latest`, with safe authentication. Cursor itself has no microphone, so the skill wires the app or a sample client rather than the IDE.

When should I use Grok Realtime Voice Integration?

Grok Realtime Voice Integration fits situations like: adding Grok speech-to-speech voice mode to an existing app; replacing an STT-LLM-TTS cascade with a single realtime voice loop; replacing an existing OpenAI Realtime voice integration with Grok.

How do I install Grok Realtime Voice Integration in Claude Code?

Run `npx skills add cursor/plugins --skill add-voice -a claude-code`. Or copy the skill folder (grok-voice/skills/add-voice in cursor/plugins) into .claude/skills/add-voice in your project. Claude Code loads it when a task matches its description.

How do I install Grok Realtime Voice Integration in Codex?

Run `npx skills add cursor/plugins --skill add-voice -a codex`. Or copy the skill folder (grok-voice/skills/add-voice in cursor/plugins) into .agents/skills/add-voice in your project. Codex loads it when a task matches its description.

Can I use Grok Realtime Voice Integration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cursor/plugins --skill add-voice -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-voice, .gemini/skills/add-voice, .github/skills/add-voice and .opencode/skills/add-voice in your project.

What does Grok Realtime Voice Integration need to run?

Going by SKILL.md and its folder, Grok Realtime Voice Integration needs credentials named XAI_API_KEY. Our summary lists: An XAI_API_KEY for server-side auth; A backend endpoint to mint ephemeral tokens for browser or mobile clients.

Does Grok Realtime Voice Integration access the network?

SKILL.md names 2 domains. In commands or code: api.x.ai; the agent is likely to contact it when it follows the instructions. As links in the text: docs.x.ai. This is read from the text; nothing was executed.

Is Grok Realtime Voice Integration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Grok Realtime Voice Integration use?

No licence was found for Grok Realtime Voice Integration or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Grok Realtime Voice Integration use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Grok Realtime Voice Integration?

Skills that share tags, products or a category with Grok Realtime Voice Integration: Deepgram Python Text-to-Speech (deepgram/deepgram-python-sdk, 469 stars), Speech Engine (elevenlabs/skills, 479 stars), Local AI Use (amd/skills, 395 stars) and Voice AI Development (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grok Realtime Voice Integration?

cursor (a GitHub organization, an official publisher) maintains it in cursor/plugins, which has 10,130 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 6, 2026.

Source: cursor/plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.