Agent skill

Speech Engine

by elevenlabs in elevenlabs/skills

Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

MITAuto-check: warningsMedia & Creative

Install Speech Engine

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add elevenlabs/skills --skill speech-engine -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install elevenlabs/skills speech-engine --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/elevenlabs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/speech-engine .claude/skills/speech-engine && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speech-engine
GitHub stars
481
Token cost
~2.5k tokens
SKILL.md length
897 words
Files
4 (incl. references)
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

  • Works in 5 steps: The browser sends user audio to… → ElevenLabs sends speech-recognition… → Your server derives trusted application… → …
  • Building Speech Engine servers
  • SKILL.md covers When to Use, How It Works, Implementation Flow and Create a Speech Engine, plus 3 more sections
  • Calls ngrok; needs ELEVENLABS_API_KEY

What it does

Speech Engine is an agent skill from elevenlabs/skills. Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. Use when building Speech Engine servers, WebSocket handlers, WebRTC browser clients, conversation token endpoints, interruption-aware streaming responses, or voice-enabled chat agents that connect developer-owned server logic to ElevenLabs speech-to-text and text-to-speech.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/installation.md`, `references/javascript-sdk-reference.md` and `references/python-sdk-reference.md`). Compatibility notes: Requires internet access and an ElevenLabs API key (ELEVENLABSAPIKEY).

It sits in Media & Creative, covering Text to speech and voice, Realtime and WebSockets and Speech recognition and synthesis. It works with ElevenLabs, Python and JavaScript. The repository describes itself as: Collections of skills for building with ElevenLabs. The licence is MIT.

When your agent uses it

  • Building Speech Engine servers
  • WebSocket handlers
  • WebRTC browser clients
  • Conversation token endpoints

Example prompts

  • “/speech-engine”

Requirements

  • Python 3
  • A credential in ELEVENLABS_API_KEY
  • Compatibility (from SKILL.md): Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. The browser sends user audio to ElevenLabs.
  2. ElevenLabs sends speech-recognition events to your server.
  3. Your server derives trusted application state without letting raw speech text control tools or privileged actions.
  4. Your server streams text back through the SDK.
  5. ElevenLabs converts the response to speech and plays it in the browser.

What it can do on your machine

Read from SKILL.md and the folder at commit 1d08a4a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ngrok

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • elevenlabs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELEVENLABS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).

    From compatibility in the SKILL.md frontmatter.

Context cost

Speech Engine loads about 2.5k tokens when it runs, and up to ~5.4k if it reads all its reference files. Until then it costs about 95 tokens; SKILL.md has 897 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningMentions a paste, webhook or tunnelling service often used to send data outSKILL.md:97
    WS_URL` should look like `wss://example.ngrok.app/ws` locally or your production WebSocket route in deployment.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from elevenlabs/skills at commit 1d08a4a, republished under its MIT licence (© elevenlabs). 897 words, ~2,507 tokens.

Download SKILL.mdSave it as .claude/skills/speech-engine/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
speech-engine
description
Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. Use when building Speech Engine servers, WebSocket handlers, WebRTC browser clients, conversation token endpoints, interruption-aware streaming responses, or voice-enabled chat agents that connect developer-owned server logic to ElevenLabs speech-to-text and text-to-speech.
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
license
MIT

ElevenLabs Speech Engine

Add a real-time voice interface to a custom agent. ElevenLabs handles microphone audio, speech-to-text, turn-taking, text-to-speech, and browser playback; your server exposes a Speech Engine WebSocket endpoint and streams response text back.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only. For deeper SDK details, read JavaScript SDK Reference or Python SDK Reference.

When to Use

Use Speech Engine when the user wants to:

  • Add voice to an existing chat app or custom server pipeline
  • Add voice to OpenClaw, Hermes, or a similar agent runtime while keeping agent logic on the developer-owned server
  • Build a developer-hosted WebSocket server for ElevenLabs voice conversations
  • Stream response text back as spoken audio after your server validates user intent
  • Handle user interruptions while a response is still streaming
  • Build a browser client with @elevenlabs/react or @elevenlabs/client using a server-issued conversation token

Use the agents skill instead when the user is creating or configuring a hosted ElevenLabs Conversational AI agent with platform-managed prompts, tools, workflows, phone numbers, or widgets.

How It Works

Each Speech Engine WebSocket connection represents one conversation.

  1. The browser sends user audio to ElevenLabs.
  2. ElevenLabs sends speech-recognition events to your server.
  3. Your server derives trusted application state without letting raw speech text control tools or privileged actions.
  4. Your server streams text back through the SDK.
  5. ElevenLabs converts the response to speech and plays it in the browser.

The SDK manages WebSocket routing, request verification, session lifecycle, ping/pong, turn-taking, and interruption handling. sendResponse() / send_response() accepts a string or async iterable of response text.

Treat speech-recognition text as untrusted user input. Do not map raw speech text directly into model roles, responses, or tool calls. Use deterministic validation, allowlisted intents, or explicit user confirmation before any transcript-derived value affects downstream response or tool logic.

Implementation Flow

  1. Install server dependencies and configure ELEVENLABS_API_KEY.
  2. Expose your Speech Engine server through a public HTTPS URL for local development, for example with ngrok http 3001.
  3. Create a Speech Engine resource with ws_url / wsUrl pointing at the public WebSocket URL, usually wss://.../ws.
  4. Store the returned Speech Engine ID, for example in ELEVENLABS_SPEECH_ENGINE_ID.
  5. Start a Speech Engine server with engine.serve(...) in Python or speechEngine.attach(...) in TypeScript.
  6. Issue browser conversation tokens from a server endpoint. Never put ELEVENLABS_API_KEY in browser code.
  7. Start the client session with conversationToken; if the agent should greet first, enable the first-message override on the Speech Engine resource, then set overrides.agent.firstMessage in the client.

Create a Speech Engine

Python
python
import asyncio
import os

from dotenv import load_dotenv
from elevenlabs import AsyncElevenLabs

load_dotenv()

elevenlabs = AsyncElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))

async def main():
    engine = await elevenlabs.speech_engine.create(
        name="My Speech Engine",
        speech_engine={"ws_url": os.environ["PUBLIC_WS_URL"]},
        overrides={"first_message": True},
    )
    print(engine.engine_id)

asyncio.run(main())
TypeScript
typescript
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";

const elevenlabs = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});

const engine = await elevenlabs.speechEngine.create({
  name: "My Speech Engine",
  speechEngine: { wsUrl: process.env.PUBLIC_WS_URL! },
  overrides: { firstMessage: true },
});

console.log(engine.engineId);

PUBLIC_WS_URL should look like wss://example.ngrok.app/ws locally or your production WebSocket route in deployment.

The create request can also configure tts, asr, turn, speech_engine.request_headers / speechEngine.requestHeaders, overrides, and privacy for custom voices, transcription keywords, turn-taking, server auth headers, client-provided first messages, and recording behavior. See the SDK reference files for expanded examples.

Server Pattern

Run the Speech Engine server at the ws_url / wsUrl configured on the resource. Keep response generation behind your own validation boundary: raw speech-recognition text should not directly control responses, tools, secrets, or other privileged actions.

Python
python
engine = await elevenlabs.speech_engine.get(os.environ["ELEVENLABS_SPEECH_ENGINE_ID"])
await engine.serve(port=3001, path="/ws", debug=True, callbacks=validated_callbacks)
Show full SKILL.md (390 more words)Show less
TypeScript
typescript
const engine = await elevenlabs.speechEngine.get(process.env.ELEVENLABS_SPEECH_ENGINE_ID!);
engine.attach(httpServer, "/ws", { debug: true, ...validatedCallbacks });

In TypeScript, pass interruption signals to downstream async work when it supports cancellation so interrupted responses stop quickly. In Python, the SDK cancels the previous turn handler when a newer turn arrives.

Server callbacks can distinguish clean closes from dropped connections: use onClose / on_close for clean disconnects and onDisconnect / on_disconnect for unexpected WebSocket drops.

Security note: speech-recognition text can contain prompt-injection attempts from user speech or played audio. Treat it as untrusted input. Convert it into trusted application state before invoking response generation, tools, or privileged workflows.

Disabling authentication (advanced, dangerous)

Both engine.attach() (TypeScript) and engine.serve() / SpeechEngineServer (Python) verify a JWT on every incoming WebSocket by default. This is what proves the connection is really coming from ElevenLabs and not from an attacker who guessed the URL. Do not turn this off.

An escape hatch exists — disableAuth: true in the callback options (TypeScript) or disable_auth=True on serve() / SpeechEngineServer(...) (Python) — for the narrow case where a compensating network-level control is already in place. Without such a control, disabling auth means any client on the internet that finds your URL can open sessions. Concretely, an attacker can:

  • open unlimited conversations to drain your ElevenLabs quota and downstream LLM budget
  • feed crafted transcripts to your response pipeline, effectively impersonating a user
  • use your server as an oracle to probe backend state, tools, or prompts

Only recommend disableAuth / disable_auth when the user has already implemented at least one of:

  • IP allowlist — the server (or an upstream firewall / load balancer / API gateway) only accepts inbound traffic from ElevenLabs' documented egress ranges.
  • Custom shared-secret header — a secret header configured on the Speech Engine resource via speech_engine.request_headers / speechEngine.requestHeaders at create time, validated by an upstream proxy (or by the developer's own middleware in front of attach() / serve()) before requests reach the SDK.

If the user cannot confirm one of the above is in place, leave the default authentication on. Skipping JWT verification without a mitigation is not an optimization or a convenience — it is unauthenticated public compute.

Browser Client

Create a server-side token endpoint and have the browser request a token before starting the microphone session. Keep the Speech Engine ID and API key on the server. If the client passes overrides.agent.firstMessage, the Speech Engine resource must have the first-message override enabled.

typescript
import express from "express";
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import "dotenv/config";

const app = express();
const elevenlabs = new ElevenLabsClient();

app.get("/api/token", async (_req, res) => {
  const response = await elevenlabs.conversationalAi.conversations.getWebrtcToken({
    agentId: process.env.ELEVENLABS_SPEECH_ENGINE_ID!,
  });
  res.json({ token: response.token });
});

React clients can use @elevenlabs/react:

tsx
import { useConversation } from "@elevenlabs/react";

export function VoiceControls() {
  const conversation = useConversation({
    onConnect: () => console.log("connected"),
    onDisconnect: () => console.log("disconnected"),
    onError: (error) => console.error(error),
  });

  async function startConversation() {
    await navigator.mediaDevices.getUserMedia({ audio: true });
    const { token } = await fetch("/api/token").then((res) => res.json());

    await conversation.startSession({
      conversationToken: token,
      overrides: {
        agent: { firstMessage: "Hello! How can I help you today?" },
      },
    });
  }

  return <button onClick={startConversation}>Start conversation</button>;
}

References

© elevenlabs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in speech-engine of elevenlabs/skills.

  • SKILL.md
  • references/installation.md
  • references/javascript-sdk-reference.md
  • references/python-sdk-reference.md

Open the folder on GitHubat commit 1d08a4a

Compare with similar skills

Speech Engine next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speech Engine compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speech Engine this skillelevenlabs/skills481—~2.5kAutomated safety check: WarnMIT
Speech To Texttadaspetra/loop2963 repos~2kAutomated safety check: PassMIT
Agentstadaspetra/loop2961 repos~2.5kAutomated safety check: PassMIT
Deepgram JS Audio Intelligencedeepgram/deepgram-js-sdk276—~1.5kAutomated safety check: PassMIT
Deepgram Python Text-to-Speechdeepgram/deepgram-python-sdk469—~1.8kAutomated safety check: PassMIT
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • Speech To Text

    tadaspetra/loop

    Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 3 repos~2k tokens
    Media & CreativeAuto-check passed
  • Agents

    tadaspetra/loop

    Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 1 repo~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Deepgram JS Audio Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…

    276 GitHub stars~1.5k tokensUpdated today
    Media & CreativeAuto-check passed
  • Deepgram Python Text-to-Speech

    deepgram/deepgram-python-sdk

    Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper.

    469 GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated yesterday
    Backend & APIsAuto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes

More from elevenlabs/skills

All 8 skills in this repo
  • Dubbing

    elevenlabs/skills

    Dub audio and video into other languages using the ElevenLabs Dubbing API (dubbingv2), preserving the original speakers' voices.

    481 GitHub stars~3.3k tokensUpdated today
    Auto-check passed
  • Text To Speech

    elevenlabs/skills

    Convert text to speech using ElevenLabs voice AI. An agent skill from elevenlabs/skills.

    481 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Voice Changer

    elevenlabs/skills

    Transform the voice in an audio recording into a different target voice while preserving emotion, timing, and delivery using the ElevenLabs Voice Changer (speech-to-speech) API.

    481 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Voice Isolator

    elevenlabs/skills

    Remove background noise and isolate vocals/speech from audio using ElevenLabs Voice Isolator (audio isolation) API.

    481 GitHub stars~923 tokensUpdated today
    Auto-check passed
  • Agents

    elevenlabs/skills

    Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills.

    481 GitHub stars~6.5k tokensUpdated today
    Auto-check passed
  • Setup API Key

    elevenlabs/skills

    Guides users through setting up an ElevenLabs API key for REST API and SDK workflows.

    481 GitHub stars~954 tokensUpdated today
    Auto-check: notes

Questions about Speech Engine

What does Speech Engine do?

Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine. Speech Engine is an agent skill from elevenlabs/skills. Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

When should I use Speech Engine?

Speech Engine fits situations like: building Speech Engine servers; webSocket handlers; webRTC browser clients; conversation token endpoints.

How do I install Speech Engine in Claude Code?

Run `npx skills add elevenlabs/skills --skill speech-engine -a claude-code`. Or copy the skill folder (speech-engine in elevenlabs/skills) into .claude/skills/speech-engine in your project. Claude Code loads it when a task matches its description.

How do I install Speech Engine in Codex?

Run `npx skills add elevenlabs/skills --skill speech-engine -a codex`. Or copy the skill folder (speech-engine in elevenlabs/skills) into .agents/skills/speech-engine in your project. Codex loads it when a task matches its description.

Can I use Speech Engine in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add elevenlabs/skills --skill speech-engine -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-engine, .gemini/skills/speech-engine, .github/skills/speech-engine and .opencode/skills/speech-engine in your project.

What does Speech Engine need to run?

Going by SKILL.md and its folder, Speech Engine needs the command-line tools its instructions call (ngrok) and credentials named ELEVENLABS_API_KEY. Our summary lists: Python 3; A credential in ELEVENLABS_API_KEY. Compatibility (from SKILL.md): Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY)..

Does Speech Engine access the network?

SKILL.md names 1 domain. As links in the text: elevenlabs.io. This is read from the text; nothing was executed.

Is Speech Engine safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): mentions a paste, webhook or tunnelling service often used to send data out. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Speech Engine use?

Speech Engine is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Speech Engine use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.9k tokens, read only when the agent opens those files.

What are the alternatives to Speech Engine?

Skills that share tags, products or a category with Speech Engine: Speech To Text (tadaspetra/loop, 296 stars), Agents (tadaspetra/loop, 296 stars), Deepgram JS Audio Intelligence (deepgram/deepgram-js-sdk, 276 stars) and Deepgram Python Text-to-Speech (deepgram/deepgram-python-sdk, 469 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speech Engine?

elevenlabs (a GitHub organization) maintains it in elevenlabs/skills, which has 481 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on October 7, 2026.

Source: elevenlabs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.