Speech To Text
tadaspetra/loop
Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.
Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add elevenlabs/skills --skill speech-engine -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install elevenlabs/skills speech-engine --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/elevenlabs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/speech-engine .claude/skills/speech-engine && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "speech-engine" agent skill from https://github.com/elevenlabs/skills/tree/main/speech-engine into .claude/skills/speech-engine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-engine", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/elevenlabs/skills/tree/main/speech-engineType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add elevenlabs/skills --skill speech-engine -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install elevenlabs/skills speech-engine --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elevenlabs/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/speech-engine .agents/skills/speech-engine && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "speech-engine" agent skill from https://github.com/elevenlabs/skills/tree/main/speech-engine into .agents/skills/speech-engine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-engine", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add elevenlabs/skills --skill speech-engine -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install elevenlabs/skills speech-engine --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elevenlabs/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/speech-engine .cursor/skills/speech-engine && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "speech-engine" agent skill from https://github.com/elevenlabs/skills/tree/main/speech-engine into .cursor/skills/speech-engine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-engine", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/elevenlabs/skills.git --path speech-engine--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add elevenlabs/skills --skill speech-engine -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install elevenlabs/skills speech-engine --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elevenlabs/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/speech-engine .gemini/skills/speech-engine && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "speech-engine" agent skill from https://github.com/elevenlabs/skills/tree/main/speech-engine into .gemini/skills/speech-engine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-engine", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install elevenlabs/skills speech-engineInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add elevenlabs/skills --skill speech-engine -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/elevenlabs/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/speech-engine .github/skills/speech-engine && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "speech-engine" agent skill from https://github.com/elevenlabs/skills/tree/main/speech-engine into .github/skills/speech-engine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-engine", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add elevenlabs/skills --skill speech-engine -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install elevenlabs/skills speech-engine --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/elevenlabs/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/speech-engine .opencode/skills/speech-engine && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "speech-engine" agent skill from https://github.com/elevenlabs/skills/tree/main/speech-engine into .opencode/skills/speech-engine/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-engine", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
speech-engineAdd real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.
Speech Engine is an agent skill from elevenlabs/skills. Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. Use when building Speech Engine servers, WebSocket handlers, WebRTC browser clients, conversation token endpoints, interruption-aware streaming responses, or voice-enabled chat agents that connect developer-owned server logic to ElevenLabs speech-to-text and text-to-speech.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/installation.md`, `references/javascript-sdk-reference.md` and `references/python-sdk-reference.md`). Compatibility notes: Requires internet access and an ElevenLabs API key (ELEVENLABSAPIKEY).
It sits in Media & Creative, covering Text to speech and voice, Realtime and WebSockets and Speech recognition and synthesis. It works with ElevenLabs, Python and JavaScript. The repository describes itself as: Collections of skills for building with ElevenLabs. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1d08a4a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ngrokFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
elevenlabs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
ELEVENLABS_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
From compatibility in the SKILL.md frontmatter.
Speech Engine loads about 2.5k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 95 tokens; SKILL.md has 897 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
WS_URL` should look like `wss://example.ngrok.app/ws` locally or your production WebSocket route in deployment.Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from elevenlabs/skills at commit 1d08a4a, republished under its MIT licence (© elevenlabs). 897 words, ~2,507 tokens.
.claude/skills/speech-engine/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Add a real-time voice interface to a custom agent. ElevenLabs handles microphone audio, speech-to-text, turn-taking, text-to-speech, and browser playback; your server exposes a Speech Engine WebSocket endpoint and streams response text back.
Setup: See Installation Guide. For JavaScript, use
@elevenlabs/*packages only. For deeper SDK details, read JavaScript SDK Reference or Python SDK Reference.
Use Speech Engine when the user wants to:
@elevenlabs/react or @elevenlabs/client using a server-issued conversation tokenUse the agents skill instead when the user is creating or configuring a hosted ElevenLabs Conversational AI agent with platform-managed prompts, tools, workflows, phone numbers, or widgets.
Each Speech Engine WebSocket connection represents one conversation.
The SDK manages WebSocket routing, request verification, session lifecycle, ping/pong, turn-taking, and interruption handling. sendResponse() / send_response() accepts a string or async iterable of response text.
Treat speech-recognition text as untrusted user input. Do not map raw speech text directly into model roles, responses, or tool calls. Use deterministic validation, allowlisted intents, or explicit user confirmation before any transcript-derived value affects downstream response or tool logic.
ELEVENLABS_API_KEY.ngrok http 3001.ws_url / wsUrl pointing at the public WebSocket URL, usually wss://.../ws.ELEVENLABS_SPEECH_ENGINE_ID.engine.serve(...) in Python or speechEngine.attach(...) in TypeScript.ELEVENLABS_API_KEY in browser code.conversationToken; if the agent should greet first, enable the first-message override on the Speech Engine resource, then set overrides.agent.firstMessage in the client.import asyncio
import os
from dotenv import load_dotenv
from elevenlabs import AsyncElevenLabs
load_dotenv()
elevenlabs = AsyncElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
async def main():
engine = await elevenlabs.speech_engine.create(
name="My Speech Engine",
speech_engine={"ws_url": os.environ["PUBLIC_WS_URL"]},
overrides={"first_message": True},
)
print(engine.engine_id)
asyncio.run(main())import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";
const elevenlabs = new ElevenLabsClient({
apiKey: process.env.ELEVENLABS_API_KEY,
});
const engine = await elevenlabs.speechEngine.create({
name: "My Speech Engine",
speechEngine: { wsUrl: process.env.PUBLIC_WS_URL! },
overrides: { firstMessage: true },
});
console.log(engine.engineId);PUBLIC_WS_URL should look like wss://example.ngrok.app/ws locally or your production WebSocket route in deployment.
The create request can also configure tts, asr, turn, speech_engine.request_headers / speechEngine.requestHeaders, overrides, and privacy for custom voices, transcription keywords, turn-taking, server auth headers, client-provided first messages, and recording behavior. See the SDK reference files for expanded examples.
Run the Speech Engine server at the ws_url / wsUrl configured on the resource. Keep response generation behind your own validation boundary: raw speech-recognition text should not directly control responses, tools, secrets, or other privileged actions.
engine = await elevenlabs.speech_engine.get(os.environ["ELEVENLABS_SPEECH_ENGINE_ID"])
await engine.serve(port=3001, path="/ws", debug=True, callbacks=validated_callbacks)const engine = await elevenlabs.speechEngine.get(process.env.ELEVENLABS_SPEECH_ENGINE_ID!);
engine.attach(httpServer, "/ws", { debug: true, ...validatedCallbacks });In TypeScript, pass interruption signals to downstream async work when it supports cancellation so interrupted responses stop quickly. In Python, the SDK cancels the previous turn handler when a newer turn arrives.
Server callbacks can distinguish clean closes from dropped connections: use onClose / on_close for clean disconnects and onDisconnect / on_disconnect for unexpected WebSocket drops.
Security note: speech-recognition text can contain prompt-injection attempts from user speech or played audio. Treat it as untrusted input. Convert it into trusted application state before invoking response generation, tools, or privileged workflows.
Both engine.attach() (TypeScript) and engine.serve() / SpeechEngineServer (Python) verify a JWT on every incoming WebSocket by default. This is what proves the connection is really coming from ElevenLabs and not from an attacker who guessed the URL. Do not turn this off.
An escape hatch exists — disableAuth: true in the callback options (TypeScript) or disable_auth=True on serve() / SpeechEngineServer(...) (Python) — for the narrow case where a compensating network-level control is already in place. Without such a control, disabling auth means any client on the internet that finds your URL can open sessions. Concretely, an attacker can:
Only recommend disableAuth / disable_auth when the user has already implemented at least one of:
speech_engine.request_headers / speechEngine.requestHeaders at create time, validated by an upstream proxy (or by the developer's own middleware in front of attach() / serve()) before requests reach the SDK.If the user cannot confirm one of the above is in place, leave the default authentication on. Skipping JWT verification without a mitigation is not an optimization or a convenience — it is unauthenticated public compute.
Create a server-side token endpoint and have the browser request a token before starting the microphone session. Keep the Speech Engine ID and API key on the server. If the client passes overrides.agent.firstMessage, the Speech Engine resource must have the first-message override enabled.
import express from "express";
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";
const app = express();
const elevenlabs = new ElevenLabsClient();
app.get("/api/token", async (_req, res) => {
const response = await elevenlabs.conversationalAi.conversations.getWebrtcToken({
agentId: process.env.ELEVENLABS_SPEECH_ENGINE_ID!,
});
res.json({ token: response.token });
});React clients can use @elevenlabs/react:
import { useConversation } from "@elevenlabs/react";
export function VoiceControls() {
const conversation = useConversation({
onConnect: () => console.log("connected"),
onDisconnect: () => console.log("disconnected"),
onError: (error) => console.error(error),
});
async function startConversation() {
await navigator.mediaDevices.getUserMedia({ audio: true });
const { token } = await fetch("/api/token").then((res) => res.json());
await conversation.startSession({
conversationToken: token,
overrides: {
agent: { firstMessage: "Hello! How can I help you today?" },
},
});
}
return <button onClick={startConversation}>Start conversation</button>;
}© elevenlabs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in speech-engine of elevenlabs/skills.
Open the folder on GitHubat commit 1d08a4a
Speech Engine next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Speech Engine this skillelevenlabs/skills | 481 | — | ~2.5k | Automated safety check: Warn | MIT | |
| Speech To Texttadaspetra/loop | 296 | 3 repos | ~2k | Automated safety check: Pass | MIT | |
| Agentstadaspetra/loop | 296 | 1 repos | ~2.5k | Automated safety check: Pass | MIT | |
| Deepgram JS Audio Intelligencedeepgram/deepgram-js-sdk | 276 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Deepgram Python Text-to-Speechdeepgram/deepgram-python-sdk | 469 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Gemini Live API Devgoogle-gemini/gemini-skills | 4.3k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 |
tadaspetra/loop
Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.
tadaspetra/loop
Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.
deepgram/deepgram-js-sdk
A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…
deepgram/deepgram-python-sdk
Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper.
google-gemini/gemini-skills
A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.
amd/skills
Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.
elevenlabs/skills
Dub audio and video into other languages using the ElevenLabs Dubbing API (dubbingv2), preserving the original speakers' voices.
elevenlabs/skills
Convert text to speech using ElevenLabs voice AI. An agent skill from elevenlabs/skills.
elevenlabs/skills
Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer (speech-to-speech) API.
elevenlabs/skills
Remove background noise and isolate vocals/speech from audio using ElevenLabs Voice Isolator (audio isolation) API.
elevenlabs/skills
Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills.
elevenlabs/skills
Guides users through setting up an ElevenLabs API key for REST API and SDK workflows.
Works with
Categories
Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. Speech Engine is an agent skill from elevenlabs/skills. Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.
Speech Engine fits situations like: building Speech Engine servers; webSocket handlers; webRTC browser clients; conversation token endpoints.
Run `npx skills add elevenlabs/skills --skill speech-engine -a claude-code`. Or copy the skill folder (speech-engine in elevenlabs/skills) into .claude/skills/speech-engine in your project. Claude Code loads it when a task matches its description.
Run `npx skills add elevenlabs/skills --skill speech-engine -a codex`. Or copy the skill folder (speech-engine in elevenlabs/skills) into .agents/skills/speech-engine in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add elevenlabs/skills --skill speech-engine -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-engine, .gemini/skills/speech-engine, .github/skills/speech-engine and .opencode/skills/speech-engine in your project.
Going by SKILL.md and its folder, Speech Engine needs the command-line tools its instructions call (ngrok) and credentials named ELEVENLABS_API_KEY. Our summary lists: Python 3; A credential in ELEVENLABS_API_KEY. Compatibility (from SKILL.md): Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY)..
SKILL.md names 1 domain. As links in the text: elevenlabs.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): mentions a paste, webhook or tunnelling service often used to send data out. Read the flagged lines before installing; the check is not a guarantee either way.
Speech Engine is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Speech Engine: Speech To Text (tadaspetra/loop, 296 stars), Agents (tadaspetra/loop, 296 stars), Deepgram JS Audio Intelligence (deepgram/deepgram-js-sdk, 276 stars) and Deepgram Python Text-to-Speech (deepgram/deepgram-python-sdk, 469 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
elevenlabs (a GitHub organization) maintains it in elevenlabs/skills, which has 481 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.
Source: elevenlabs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.