Speech Engine
elevenlabs/skills
Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.
Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.
$ npx skills add tadaspetra/loop --skill speech-to-text -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tadaspetra/loop speech-to-text --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tadaspetra/loop.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/speech-to-text .claude/skills/speech-to-text && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "speech-to-text" agent skill from https://github.com/tadaspetra/loop/tree/main/.agents/skills/speech-to-text into .claude/skills/speech-to-text/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-to-text", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tadaspetra/loop/tree/main/.agents/skills/speech-to-textType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tadaspetra/loop --skill speech-to-text -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tadaspetra/loop speech-to-text --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tadaspetra/loop.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/speech-to-text .agents/skills/speech-to-text && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "speech-to-text" agent skill from https://github.com/tadaspetra/loop/tree/main/.agents/skills/speech-to-text into .agents/skills/speech-to-text/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-to-text", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tadaspetra/loop --skill speech-to-text -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tadaspetra/loop speech-to-text --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tadaspetra/loop.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/speech-to-text .cursor/skills/speech-to-text && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "speech-to-text" agent skill from https://github.com/tadaspetra/loop/tree/main/.agents/skills/speech-to-text into .cursor/skills/speech-to-text/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-to-text", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tadaspetra/loop.git --path .agents/skills/speech-to-text--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tadaspetra/loop --skill speech-to-text -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tadaspetra/loop speech-to-text --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tadaspetra/loop.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/speech-to-text .gemini/skills/speech-to-text && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "speech-to-text" agent skill from https://github.com/tadaspetra/loop/tree/main/.agents/skills/speech-to-text into .gemini/skills/speech-to-text/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-to-text", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tadaspetra/loop speech-to-textInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tadaspetra/loop --skill speech-to-text -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tadaspetra/loop.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/speech-to-text .github/skills/speech-to-text && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "speech-to-text" agent skill from https://github.com/tadaspetra/loop/tree/main/.agents/skills/speech-to-text into .github/skills/speech-to-text/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-to-text", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tadaspetra/loop --skill speech-to-text -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tadaspetra/loop speech-to-text --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tadaspetra/loop.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/speech-to-text .opencode/skills/speech-to-text && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "speech-to-text" agent skill from https://github.com/tadaspetra/loop/tree/main/.agents/skills/speech-to-text into .opencode/skills/speech-to-text/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "speech-to-text", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
speech-to-textTranscribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.
Speech To Text is an agent skill from tadaspetra/loop. Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/installation.md`, `references/realtime-client-side.md` and `references/realtime-commit-strategies.md`). Compatibility notes: Requires internet access and an ElevenLabs API key (ELEVENLABSAPIKEY).
It sits in Media & Creative, covering Transcription, Speech recognition and synthesis and Text to speech and voice. It works with ElevenLabs and JavaScript. The repository describes itself as: Record, Cut, Edit, Render with AI. The licence is MIT.
Read from SKILL.md and the folder at commit 452e950. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.elevenlabs.ioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
ELEVENLABS_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
From compatibility in the SKILL.md frontmatter.
Speech To Text loads about 2k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 377 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from tadaspetra/loop at commit 452e950, republished under its MIT licence (© tadaspetra). 377 words, ~2,040 tokens.
.claude/skills/speech-to-text/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.Transcribe audio to text with Scribe v2 - supports 90+ languages, speaker diarization, and word-level timestamps.
Setup: See Installation Guide. For JavaScript, use
@elevenlabs/*packages only.
from elevenlabs import ElevenLabs
client = ElevenLabs()
with open("audio.mp3", "rb") as audio_file:
result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")
print(result.text)import { ElevenLabsClient } from '@elevenlabs/elevenlabs-js';
import { createReadStream } from 'fs';
const client = new ElevenLabsClient();
const result = await client.speechToText.convert({
file: createReadStream('audio.mp3'),
modelId: 'scribe_v2'
});
console.log(result.text);curl -X POST "https://api.elevenlabs.io/v1/speech-to-text" \
-H "xi-api-key: $ELEVENLABS_API_KEY" -F "file=@audio.mp3" -F "model_id=scribe_v2"| Model ID | Description | Best For |
|---|---|---|
scribe_v2 | State-of-the-art accuracy, 90+ languages | Batch transcription, subtitles, long-form audio |
scribe_v2_realtime | Low latency (~150ms) | Live transcription, voice agents |
Word-level timestamps include type classification and speaker identification:
result = client.speech_to_text.convert(
file=audio_file, model_id="scribe_v2", timestamps_granularity="word"
)
for word in result.words:
print(f"{word.text}: {word.start}s - {word.end}s (type: {word.type})")
Identify WHO said WHAT - the model labels each word with a speaker ID, useful for meetings, interviews, or any multi-speaker audio:
result = client.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
diarize=True
)
for word in result.words:
print(f"[{word.speaker_id}] {word.text}")Help the model recognize specific words it might otherwise mishear - product names, technical jargon, or unusual spellings (up to 100 terms):
result = client.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
keyterms=["ElevenLabs", "Scribe", "API"]
)Automatic detection with optional language hint:
result = client.speech_to_text.convert(
file=audio_file,
model_id="scribe_v2",
language_code="eng" # ISO 639-1 or ISO 639-3 code
)
print(f"Detected: {result.language_code} ({result.language_probability:.0%})")Audio: MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus Video: MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP
Limits: Up to 3GB file size, 10 hours duration
{
"text": "The full transcription text",
"language_code": "eng",
"language_probability": 0.98,
"words": [
{ "text": "The", "start": 0.0, "end": 0.15, "type": "word", "speaker_id": "speaker_0" },
{ "text": " ", "start": 0.15, "end": 0.16, "type": "spacing", "speaker_id": "speaker_0" }
]
}Word types:
word - An actual spoken wordspacing - Whitespace between words (useful for precise timing)audio_event - Non-speech sounds the model detected (laughter, applause, music, etc.)try:
result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")
except Exception as e:
print(f"Transcription failed: {e}")Common errors:
Monitor usage via request-id response header:
response = client.speech_to_text.convert.with_raw_response(file=audio_file, model_id="scribe_v2")
result = response.parse()
print(f"Request ID: {response.headers.get('request-id')}")For live transcription with ultra-low latency (~150ms), use the real-time API. The real-time API produces two types of transcripts:
A "commit" tells the model to finalize the current segment. You can commit manually (e.g., when the user pauses) or use Voice Activity Detection (VAD) to auto-commit on silence.
import asyncio
from elevenlabs import ElevenLabs
client = ElevenLabs()
async def transcribe_realtime():
async with client.speech_to_text.realtime.connect(
model_id="scribe_v2_realtime",
include_timestamps=True,
) as connection:
await connection.stream_url("https://example.com/audio.mp3")
async for event in connection:
if event.type == "partial_transcript":
print(f"Partial: {event.text}")
elif event.type == "committed_transcript":
print(f"Final: {event.text}")
asyncio.run(transcribe_realtime())import { useScribe } from "@elevenlabs/react";
function TranscriptionComponent() {
const [transcript, setTranscript] = useState("");
const scribe = useScribe({
modelId: "scribe_v2_realtime",
onPartialTranscript: (data) => console.log("Partial:", data.text),
onCommittedTranscript: (data) => setTranscript((prev) => prev + data.text),
});
const start = async () => {
// Get token from your backend (never expose API key to client)
const { token } = await fetch("/scribe-token").then((r) => r.json());
await scribe.connect({
token,
microphone: { echoCancellation: true, noiseSuppression: true },
});
};
return <button onClick={start}>Start Recording</button>;
}| Strategy | Description |
|---|---|
| Manual | You call commit() when ready - use for file processing or when you control the audio segments |
| VAD | Voice Activity Detection auto-commits when silence is detected - use for live microphone input |
// VAD configuration
const connection = await client.speechToText.realtime.connect({
modelId: 'scribe_v2_realtime',
vad: {
silenceThresholdSecs: 1.5,
threshold: 0.4
}
});| Event | Description |
|---|---|
partial_transcript | Live interim results |
committed_transcript | Final results after commit |
committed_transcript_with_timestamps | Final with word timing |
error | Error occurred |
See real-time references for complete documentation.
© tadaspetra, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (references) in .agents/skills/speech-to-text of tadaspetra/loop.
Open the folder on GitHubat commit 452e950
We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in tadaspetra/loop, which our catalogue first saw on October 7, 2026.
Speech To Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Speech To Text this skilltadaspetra/loop | 296 | 2 repos | ~2k | Automated safety check: Pass | MIT | |
| Speech Engineelevenlabs/skills | 482 | — | ~2.5k | Automated safety check: Warn | MIT | |
| Local AI Useamd/skills | 408 | — | ~5k | Automated safety check: Notes | MIT | |
| Summarize Callreysu/ai-life-skills | 270 | — | ~3.8k | Automated safety check: Notes | MIT | |
| C Voicedaxaur/openpaw | 173 | — | ~401 | Automated safety check: Pass | MIT | |
| Elevenlabs Core Workflow Bjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT |
elevenlabs/skills
Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.
amd/skills
Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.
reysu/ai-life-skills
Transcribe a call recording with speaker diarization, summarize it, and create Obsidian vault notes (call note, transcript, person notes for participants).
daxaur/openpaw
Convert speech to text using sag (ElevenLabs STT) and synthesize speech using say (macOS built-in TTS).
jeremylongshore/tons-of-skills-marketplace
Implement ElevenLabs speech-to-speech, sound effects, audio isolation, and speech-to-text.
shang-zhu/violin
Dub a video into another language and generate subtitles using the default Together + Cartesia stack.
tadaspetra/loop
Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.
tadaspetra/loop
Generate sound effects from text descriptions using ElevenLabs.
tadaspetra/loop
Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.
Works with
Categories
Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop. Speech To Text is an agent skill from tadaspetra/loop. Transcribe audio to text using ElevenLabs Scribe v2.
Speech To Text fits situations like: converting audio/video to text; generating subtitles; transcribing meetings; processing spoken content.
Run `npx skills add tadaspetra/loop --skill speech-to-text -a claude-code`. Or copy the skill folder (.agents/skills/speech-to-text in tadaspetra/loop) into .claude/skills/speech-to-text in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tadaspetra/loop --skill speech-to-text -a codex`. Or copy the skill folder (.agents/skills/speech-to-text in tadaspetra/loop) into .agents/skills/speech-to-text in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tadaspetra/loop --skill speech-to-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-to-text, .gemini/skills/speech-to-text, .github/skills/speech-to-text and .opencode/skills/speech-to-text in your project.
Going by SKILL.md and its folder, Speech To Text needs the command-line tools its instructions call (curl) and credentials named ELEVENLABS_API_KEY. Our summary lists: Python 3; A credential in ELEVENLABS_API_KEY. Compatibility (from SKILL.md): Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY)..
SKILL.md names 1 domain. In commands or code: api.elevenlabs.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Speech To Text is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Speech To Text: Speech Engine (elevenlabs/skills, 482 stars), Local AI Use (amd/skills, 408 stars), Summarize Call (reysu/ai-life-skills, 270 stars) and C Voice (daxaur/openpaw, 173 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tadaspetra (a GitHub user) maintains it in tadaspetra/loop, which has 296 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 2, 2026.
Source: tadaspetra/loop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.