Agent skill

Elevenlabs Performance Tuning

by jeremylongshore in jeremylongshore/tons-of-skills-marketplace

Optimize ElevenLabs TTS latency with model selection, streaming, caching, and audio format tuning.

MITAuto-check passedMedia & Creative

Install Elevenlabs Performance Tuning

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill elevenlabs-performance-tuning -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace elevenlabs-performance-tuning --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/elevenlabs-performance-tuning .claude/skills/elevenlabs-performance-tuning && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
elevenlabs-performance-tuning
GitHub stars
2.8k
Token cost
~2k tokens
SKILL.md length
680 words
Files
2 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Optimize ElevenLabs TTS latency with model selection, streaming, caching, and audio format tuning.

  • Works in 6 steps: Model Selection for Latency → Output Format Optimization → HTTP Streaming for Time-to-First-Byte → …
  • Experiencing slow TTS responses
  • SKILL.md covers Overview, Prerequisites, Instructions and Output, plus 5 more sections
  • Needs ELEVENLABS_API_KEY

What it does

Elevenlabs Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize ElevenLabs TTS latency with model selection, streaming, caching, and audio format tuning. Use when experiencing slow TTS responses, implementing real-time voice features, or optimizing audio generation throughput. Trigger with "elevenlabs performance", "optimize elevenlabs", "elevenlabs latency", "elevenlabs slow", "fast TTS", "reduce elevenlabs latency", or "TTS streaming".

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/implementation.md`). Compatibility notes: Designed for Claude Code

It sits in Media & Creative, covering Text to speech and voice. It works with ElevenLabs. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Experiencing slow TTS responses
  • Implementing real-time voice features
  • Optimizing audio generation throughput
  • With elevenlabs performance

Example prompts

  • “elevenlabs performance”
  • “optimize elevenlabs”
  • “elevenlabs latency”
  • “/elevenlabs-performance-tuning”

Requirements

  • A credential in ELEVENLABS_API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code
  • Pre-approved tools (allowed-tools): Read, Write, Edit

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Model Selection for Latency
  2. Output Format Optimization
  3. HTTP Streaming for Time-to-First-Byte
  4. WebSocket Streaming for Lowest Latency
  5. Audio Caching
  6. Parallel Generation

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • elevenlabs.io
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELEVENLABS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code

    From compatibility in the SKILL.md frontmatter.

Context cost

Elevenlabs Performance Tuning loads about 2k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 680 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 680 words, ~2,021 tokens.

Download SKILL.mdSave it as .claude/skills/elevenlabs-performance-tuning/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
elevenlabs-performance-tuning
description
Optimize ElevenLabs TTS latency with model selection, streaming, caching, and audio format tuning. Use when experiencing slow TTS responses, implementing real-time voice features, or optimizing audio generation throughput. Trigger with "elevenlabs performance", "optimize elevenlabs", "elevenlabs latency", "elevenlabs slow", "fast TTS", "reduce elevenlabs latency", or "TTS streaming".
allowed-tools
Read, Write, Edit
compatibility
Designed for Claude Code
version
1.6.0
license
MIT
author
Jeremy Longshore <jeremy@intentsolutions.io>
tags
saas, voice, ai, elevenlabs, performance, optimization

ElevenLabs Performance Tuning

Overview

Optimize ElevenLabs TTS latency and throughput through model selection, streaming strategies, audio format tuning, and caching. Latency ranges from ~75ms (Flash) to ~500ms (v3) depending on configuration.

The two highest-leverage, lowest-effort levers — model choice (Step 1) and output format (Step 2) — are documented inline below. The four deeper integrations (HTTP streaming, WebSocket streaming, caching, parallel generation) are summarized here with copy-ready code in the full implementation walkthrough.

Prerequisites

  • ElevenLabs SDK installed (@elevenlabs/elevenlabs-js)
  • An ElevenLabs API key exported as ELEVENLABS_API_KEY (used by the SDK and passed as xi_api_key on the WebSocket handshake)
  • Understanding of your latency requirements
  • Audio playback infrastructure (browser, mobile, server-side)

Instructions

Step 1: Model Selection for Latency

The single biggest performance lever is model choice:

ModelAvg LatencyQualityLanguagesUse Case
eleven_flash_v2_5~75msGood32Real-time chat, IVR, gaming
eleven_turbo_v2_5~150msGood32Balanced speed/quality
eleven_multilingual_v2~300msHigh29Narration, content creation
eleven_v3~500msHighest70+Maximum expressiveness
typescript
// Select model based on use case
function selectModel(useCase: "realtime" | "balanced" | "quality" | "max_quality"): string {
  const models = {
    realtime:    "eleven_flash_v2_5",
    balanced:    "eleven_turbo_v2_5",
    quality:     "eleven_multilingual_v2",
    max_quality: "eleven_v3",
  };
  return models[useCase];
}
Step 2: Output Format Optimization

Smaller formats = faster transfer:

FormatSize/SecondQualityBest For
mp3_44100_128~16 KB/sHighDownloads, archival
mp3_22050_32~4 KB/sMediumStreaming, mobile
pcm_16000~32 KB/sRawServer-side processing
pcm_44100~88 KB/sRawHigh-quality processing
ulaw_8000~8 KB/sPhoneTelephony/IVR
typescript
// Use smaller format for streaming, higher quality for downloads
const streamingConfig = {
  output_format: "mp3_22050_32",  // 4 KB/s — fast streaming
  model_id: "eleven_flash_v2_5",   // ~75ms first byte
};

const downloadConfig = {
  output_format: "mp3_44100_128", // 16 KB/s — high quality
  model_id: "eleven_multilingual_v2",
};
Step 3: HTTP Streaming for Time-to-First-Byte

Call client.textToSpeech.stream() instead of .convert() and write each chunk to the response as it arrives, so playback starts before generation finishes — roughly halving time-to-first-byte. Set style: 0.0 in voice_settings to shave another 10–20%. Full server handler: implementation.md § Step 3.

Step 4: WebSocket Streaming for Lowest Latency

For interactive apps where text arrives incrementally (e.g., an LLM token stream), open a stream-input WebSocket, sendText() chunks as they arrive, and tune chunk_length_schedule — fewer characters per chunk means lower latency but less prosody context. Full bidirectional client: implementation.md § Step 4.

Step 5: Audio Caching

Cache generated audio for repeated content (greetings, prompts, errors) in an LRU cache keyed by a SHA-256 of voiceId:modelId:text, so a changed voice or model never serves stale audio. This eliminates ~99% of latency for repeated phrases. Full cachedTTS helper: implementation.md § Step 5.

Step 6: Parallel Generation

Generate multiple segments concurrently with a p-queue whose concurrency matches your plan's request limit (going higher returns 429s, not more throughput). Full chapter-generator: implementation.md § Step 6.

Show full SKILL.md (312 more words)Show less

Output

Applying these levers produces:

  • A model + output-format choice matched to the use case (Steps 1–2).
  • A streaming code path (HTTP or WebSocket) that logs measured time-to-first-byte, e.g. Time to first byte: 78ms / WebSocket TTFB: 91ms.
  • An LRU audio cache emitting [Cache HIT] / [Cache MISS] telemetry for repeated content.
  • A concurrency-bounded batch path that logs per-segment generation time.

Expected latency after tuning: ~75–150ms first byte on Flash/Turbo with streaming, versus ~300–500ms for a blocking convert() call on a higher-quality model.

Performance Optimization Checklist

OptimizationLatency ImpactImplementation
Flash model-60% vs v2, -85% vs v3Change model_id
Streaming endpoint-50% time-to-first-byteUse .stream() instead of .convert()
WebSocket streamingBest for LLM integrationSee Step 4
Smaller output format-30% transfer timemp3_22050_32 vs mp3_44100_128
Audio caching-99% for repeated contentLRU cache with SHA-256 keys
style: 0-10-20% latencyRemove style exaggeration
Concurrency queueMaximize throughputp-queue matching plan limit

Error Handling

IssueCauseSolution
High TTFBWrong modelSwitch to eleven_flash_v2_5
Choppy streamingNetwork bufferingUse pcm_16000 for direct playback
Cache miss stormTTL expired for popular contentUse stale-while-revalidate pattern
WebSocket dropsNetwork instabilityReconnect with buffered text
Memory pressureAudio cache too largeSet maxSize limit on LRU cache
HTTP 429Concurrency above plan limitLower p-queue concurrency

Examples

Real-time IVR (lowest latency). Pick eleven_flash_v2_5 + ulaw_8000 via selectModel("realtime"), then stream over HTTP:

typescript
await streamToResponse(greeting, voiceId, res); // logs "Time to first byte: 78ms"

LLM voice agent (incremental text). Open a WebSocket and forward tokens as they stream from the model, ending with finish():

typescript
const stream = await createTTSStream({ voiceId, chunkLengthSchedule: [50, 100, 150] });
stream.sendText("Hello, "); stream.sendText("how are you?");
const audio = await stream.finish();

Audiobook batch (throughput). Cache repeated phrases and generate chapters concurrently:

typescript
const buffers = await generateChapters(chapters, voiceId); // 5-wide, cache-backed

Full, runnable versions of every snippet above are in the implementation walkthrough.

Resources

Next Steps

For cost optimization once latency is tuned, see the elevenlabs-cost-tuning skill, which covers character-usage budgeting, model-tier cost tradeoffs, and cache-hit-rate targets.

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/.curated/elevenlabs-performance-tuning of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • references/implementation.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Elevenlabs Performance Tuning next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Elevenlabs Performance Tuning compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Elevenlabs Performance Tuning this skilljeremylongshore/tons-of-skills-marketplace2.8k—~2kAutomated safety check: PassMIT
Speech Engineelevenlabs/skills482—~2.5kAutomated safety check: WarnMIT
Setup API Keyelevenlabs/skills482—~954Automated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
ElevenLabs Voiceover Generatordigitalsamba/claude-code-video-toolkit2.2k1 repos~2.7kAutomated safety check: NotesMIT
AI SDK Developmenttrypostit/trypost6921 repos~3.5kAutomated safety check: PassMIT

Similar skills

  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • Setup API Key

    elevenlabs/skills

    Guides users through setting up an ElevenLabs API key for REST API and SDK workflows.

    482 GitHub stars~954 tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes
  • AI SDK Development

    trypostit/trypost

    TRIGGER when working with ai-sdk which is Laravel official first-party AI SDK.

    692 GitHub starsUsed in 1 repo~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Works with

Questions about Elevenlabs Performance Tuning

What does Elevenlabs Performance Tuning do?

Optimize ElevenLabs TTS latency with model selection, streaming, caching, and audio format tuning. Elevenlabs Performance Tuning is an agent skill from jeremylongshore/tons-of-skills-marketplace. Optimize ElevenLabs TTS latency with model selection, streaming, caching, and audio format tuning.

When should I use Elevenlabs Performance Tuning?

Elevenlabs Performance Tuning fits situations like: experiencing slow TTS responses; implementing real-time voice features; optimizing audio generation throughput; with elevenlabs performance.

How do I install Elevenlabs Performance Tuning in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill elevenlabs-performance-tuning -a claude-code`. Or copy the skill folder (skills/.curated/elevenlabs-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/elevenlabs-performance-tuning in your project. Claude Code loads it when a task matches its description.

How do I install Elevenlabs Performance Tuning in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill elevenlabs-performance-tuning -a codex`. Or copy the skill folder (skills/.curated/elevenlabs-performance-tuning in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/elevenlabs-performance-tuning in your project. Codex loads it when a task matches its description.

Can I use Elevenlabs Performance Tuning in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill elevenlabs-performance-tuning -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/elevenlabs-performance-tuning, .gemini/skills/elevenlabs-performance-tuning, .github/skills/elevenlabs-performance-tuning and .opencode/skills/elevenlabs-performance-tuning in your project.

What does Elevenlabs Performance Tuning need to run?

Going by SKILL.md and its folder, Elevenlabs Performance Tuning needs credentials named ELEVENLABS_API_KEY. Our summary lists: A credential in ELEVENLABS_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit. Compatibility (from SKILL.md): Designed for Claude Code.

Does Elevenlabs Performance Tuning access the network?

SKILL.md names 2 domains. As links in the text: elevenlabs.io and github.com. This is read from the text; nothing was executed.

Is Elevenlabs Performance Tuning safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Elevenlabs Performance Tuning use?

Elevenlabs Performance Tuning is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Elevenlabs Performance Tuning use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Elevenlabs Performance Tuning?

Skills that share tags, products or a category with Elevenlabs Performance Tuning: Speech Engine (elevenlabs/skills, 482 stars), Setup API Key (elevenlabs/skills, 482 stars), Music (tadaspetra/loop, 296 stars) and ElevenLabs Voiceover Generator (digitalsamba/claude-code-video-toolkit, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Elevenlabs Performance Tuning?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.