Official agent skill

Grok Dictation for Apps

by cursor in cursor/plugins

Wires xAI's Grok speech-to-text into an app you already have, covering voice input for a message box, captions shown as people talk and batch transcription of audio files.

OfficialNo licenceAuto-check passedMedia & Creative

Install Grok Dictation for Apps

skills CLI
$ npx skills add cursor/plugins --skill add-dictation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cursor/plugins add-dictation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cursor/plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/grok-voice/skills/add-dictation .claude/skills/add-dictation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-dictation
GitHub stars
10k
Token cost
~2.5k tokens
SKILL.md length
689 words
Files
1
Skills in repo
99
Repo updated
First seen
Licence
None found

At a glance

Wires xAI's Grok speech-to-text into an app you already have, covering voice input for a message box, captions shown as people talk and batch transcription of audio files.

  • Works in 4 steps: Map the app → Batch path (default) → Streaming path → …
  • Adding a mic button that dictates into a chat composer
  • SKILL.md covers Docs, Model, Pick the path and Auth, plus 2 more sections
  • Calls curl; reaches api.x.ai; needs XAI_API_KEY

What it does

The skill runs on /add-dictation and wires speech-to-text into your app rather than the IDE, since Cursor has no microphone. It tells the agent to read the model-selection section of the xAI docs before any call and pass the latest listed model instead of inventing IDs. Two paths are offered: batch, a single POST to the speech-to-text endpoint that suits a composer mic button, uploaded files and URLs and keeps the key on the server, and streaming over a WebSocket through a backend relay when users need text while they speak.

Authentication is a bearer XAI_API_KEY used server-side only, never placed in a client bundle or pasted in chat. Browsers cannot set WebSocket headers, so streaming goes through your backend. The steps start by mapping the app: the composer component, whether text inserts at the cursor or replaces, the server framework and the package manager. The mic icon belongs to dictation, and next to an /add-voice waveform button it becomes a secondary ghost button. The excerpt is cut off at the batch path.

When your agent uses it

  • Adding a mic button that dictates into a chat composer
  • Showing live captions while a user speaks
  • Transcribing uploaded audio files or URLs with word timestamps
  • Producing subtitles or meeting notes from recordings

Example prompts

  • “/add-dictation for our Next.js chat composer.”
  • “Add a mic button that transcribes a short recording and inserts the text at the cursor.”
  • “Transcribe the uploaded meeting recordings with speaker labels and word timestamps.”
  • “Set up streaming captions through a backend relay so the API key stays server-side.”

Requirements

  • An xAI API key stored server-side as XAI_API_KEY
  • A backend that can call the speech-to-text API

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Map the app
  2. Batch path (default)
  3. Streaming path
  4. Options (query params for streaming, form fields for batch)

What it can do on your machine

Read from SKILL.md and the folder at commit 9f451cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.x.ai

    Also links to:

    • docs.x.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • XAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grok Dictation for Apps loads about 2.5k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 689 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 689 words (~2,543 tokens).

“Add Grok Speech to Text to an existing app: a mic button that dictates into the composer, live captions, or transcripts of recorded audio. Run on /add-dictation, typed Dictate, or clear “transcribe” intent. Cursor has no mic; wire the app…”

— opening of SKILL.md by cursor
name
add-dictation

Read the full SKILL.md on GitHub

Files

Just SKILL.md in grok-voice/skills/add-dictation of cursor/plugins.

Open the folder on GitHubat commit 9f451cf

Compare with similar skills

Grok Dictation for Apps next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grok Dictation for Apps compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grok Dictation for Apps this skillcursor/plugins10k—~2.5kAutomated safety check: PassNone
9Router Speech-to-Textdecolua/9router30k—~745Automated safety check: PassMIT
Deepgram JS Audio Intelligencedeepgram/deepgram-js-sdk276—~1.5kAutomated safety check: PassMIT
Local AI Useamd/skills395—~5kAutomated safety check: NotesMIT
Watchmathiaschu/watch141—~4kAutomated safety check: WarnMIT
Deepgram JS Speech To Textdeepgram/deepgram-js-sdk276—~1.8kAutomated safety check: PassMIT

Similar skills

  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    30k GitHub stars~745 tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • Deepgram JS Audio Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…

    276 GitHub stars~1.5k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    395 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    141 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Deepgram JS Speech To Text

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription.

    276 GitHub stars~1.8k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    Kevin100202/Rocket-Design-and-Manufacturing-Automation-Program-from-Shenzhen22highschool

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    170 GitHub stars~518 tokensUpdated 4 mo ago
    Media & CreativeAuto-check passed

More from cursor/plugins

All 99 skills in this repo
  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 9 repos~1.6k tokens
    Auto-check passed
  • Official

    Digs into why code is shaped the way it is by checking git history, pull requests and connected tools in parallel, then reporting a cited read on the tradeoffs.

    10k GitHub starsUsed in 9 repos~2.6k tokens
    Auto-check passed
  • Official

    Starts three parallel reviewer subagents over the current conversation transcript, then turns their findings into concrete edits to existing skills.

    10k GitHub starsUsed in 5 repos~1.2k tokens
    Auto-check passed
  • Official

    Applies four layers of technical-writing rules to docs, RFCs, readmes, PR descriptions and commit messages so a tired engineer follows them on the first read.

    10k GitHub starsUsed in 10 repos~2.4k tokens
    Auto-check passed
  • Advisor Mode

    cursor/plugins

    Official

    Adds a second, stronger model that the main agent consults before major decisions, when stuck and before finishing, controlled by /advisor commands.

    10k GitHub stars~2.6k tokensUpdated today
    Auto-check: notes
  • Official

    Prepare PRs for review by cleaning noisy history, improving PR descriptions, and adding reviewer guidance without changing code behavior.

    10k GitHub starsUsed in 3 repos~569 tokens
    Auto-check passed

Works with

Questions about Grok Dictation for Apps

What does Grok Dictation for Apps do?

Wires xAI's Grok speech-to-text into an app you already have, covering voice input for a message box, captions shown as people talk and batch transcription of audio files. The skill runs on /add-dictation and wires speech-to-text into your app rather than the IDE, since Cursor has no microphone. It tells the agent to read the model-selection section of the xAI docs before any call and pass the latest listed model instead of inventing IDs.

When should I use Grok Dictation for Apps?

Grok Dictation for Apps fits situations like: adding a mic button that dictates into a chat composer; showing live captions while a user speaks; transcribing uploaded audio files or URLs with word timestamps; producing subtitles or meeting notes from recordings.

How do I install Grok Dictation for Apps in Claude Code?

Run `npx skills add cursor/plugins --skill add-dictation -a claude-code`. Or copy the skill folder (grok-voice/skills/add-dictation in cursor/plugins) into .claude/skills/add-dictation in your project. Claude Code loads it when a task matches its description.

How do I install Grok Dictation for Apps in Codex?

Run `npx skills add cursor/plugins --skill add-dictation -a codex`. Or copy the skill folder (grok-voice/skills/add-dictation in cursor/plugins) into .agents/skills/add-dictation in your project. Codex loads it when a task matches its description.

Can I use Grok Dictation for Apps in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cursor/plugins --skill add-dictation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-dictation, .gemini/skills/add-dictation, .github/skills/add-dictation and .opencode/skills/add-dictation in your project.

What does Grok Dictation for Apps need to run?

Going by SKILL.md and its folder, Grok Dictation for Apps needs the command-line tools its instructions call (curl) and credentials named XAI_API_KEY. Our summary lists: An xAI API key stored server-side as XAI_API_KEY; A backend that can call the speech-to-text API.

Does Grok Dictation for Apps access the network?

SKILL.md names 2 domains. In commands or code: api.x.ai; the agent is likely to contact it when it follows the instructions. As links in the text: docs.x.ai. This is read from the text; nothing was executed.

Is Grok Dictation for Apps safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Grok Dictation for Apps use?

No licence was found for Grok Dictation for Apps or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Grok Dictation for Apps use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Grok Dictation for Apps?

Skills that share tags, products or a category with Grok Dictation for Apps: 9Router Speech-to-Text (decolua/9router, 30k stars), Deepgram JS Audio Intelligence (deepgram/deepgram-js-sdk, 276 stars), Local AI Use (amd/skills, 395 stars) and Watch (mathiaschu/watch, 141 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grok Dictation for Apps?

cursor (a GitHub organization, an official publisher) maintains it in cursor/plugins, which has 10,130 GitHub stars. The repository holds 99 skills in this directory. The repository was last updated on October 6, 2026.

Source: cursor/plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.