Agent skill

Speech To Text

by tadaspetra in tadaspetra/loop

Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.

MITAuto-check passedMedia & Creative

Install Speech To Text

skills CLI
$ npx skills add tadaspetra/loop --skill speech-to-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tadaspetra/loop speech-to-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tadaspetra/loop.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/speech-to-text .claude/skills/speech-to-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speech-to-text
GitHub stars
296
Used in
2 other repos
Token cost
~2k tokens
SKILL.md length
377 words
Files
7 (incl. references)
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.

  • Converting audio/video to text
  • SKILL.md covers Quick Start, Models, Transcription with Timestamps and Speaker Diarization, plus 8 more sections
  • Calls curl; reaches api.elevenlabs.io; needs ELEVENLABS_API_KEY
  • Generating subtitles

What it does

Speech To Text is an agent skill from tadaspetra/loop. Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/installation.md`, `references/realtime-client-side.md` and `references/realtime-commit-strategies.md`). Compatibility notes: Requires internet access and an ElevenLabs API key (ELEVENLABSAPIKEY).

It sits in Media & Creative, covering Transcription, Speech recognition and synthesis and Text to speech and voice. It works with ElevenLabs and JavaScript. The repository describes itself as: Record, Cut, Edit, Render with AI. The licence is MIT.

When your agent uses it

  • Converting audio/video to text
  • Generating subtitles
  • Transcribing meetings
  • Processing spoken content

Example prompts

  • “/speech-to-text”

Requirements

  • Python 3
  • A credential in ELEVENLABS_API_KEY
  • Compatibility (from SKILL.md): Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).

What it can do on your machine

Read from SKILL.md and the folder at commit 452e950. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.elevenlabs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELEVENLABS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).

    From compatibility in the SKILL.md frontmatter.

Context cost

Speech To Text loads about 2k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 46 tokens; SKILL.md has 377 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~46
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tadaspetra/loop at commit 452e950, republished under its MIT licence (© tadaspetra). 377 words, ~2,040 tokens.

Download SKILL.mdSave it as .claude/skills/speech-to-text/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
speech-to-text
description
Transcribe audio to text using ElevenLabs Scribe v2. Use when converting audio/video to text, generating subtitles, transcribing meetings, or processing spoken content.
compatibility
Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY).
license
MIT

ElevenLabs Speech-to-Text

Transcribe audio to text with Scribe v2 - supports 90+ languages, speaker diarization, and word-level timestamps.

Setup: See Installation Guide. For JavaScript, use @elevenlabs/* packages only.

Quick Start

Python
python
from elevenlabs import ElevenLabs

client = ElevenLabs()

with open("audio.mp3", "rb") as audio_file:
    result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")

print(result.text)
JavaScript
javascript
import { ElevenLabsClient } from '@elevenlabs/elevenlabs-js';
import { createReadStream } from 'fs';

const client = new ElevenLabsClient();
const result = await client.speechToText.convert({
  file: createReadStream('audio.mp3'),
  modelId: 'scribe_v2'
});
console.log(result.text);
cURL
bash
curl -X POST "https://api.elevenlabs.io/v1/speech-to-text" \
  -H "xi-api-key: $ELEVENLABS_API_KEY" -F "file=@audio.mp3" -F "model_id=scribe_v2"

Models

Model IDDescriptionBest For
scribe_v2State-of-the-art accuracy, 90+ languagesBatch transcription, subtitles, long-form audio
scribe_v2_realtimeLow latency (~150ms)Live transcription, voice agents

Transcription with Timestamps

Word-level timestamps include type classification and speaker identification:

python
result = client.speech_to_text.convert(
    file=audio_file, model_id="scribe_v2", timestamps_granularity="word"
)

for word in result.words:
    print(f"{word.text}: {word.start}s - {word.end}s (type: {word.type})")

Speaker Diarization

Identify WHO said WHAT - the model labels each word with a speaker ID, useful for meetings, interviews, or any multi-speaker audio:

python
result = client.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    diarize=True
)

for word in result.words:
    print(f"[{word.speaker_id}] {word.text}")

Keyterm Prompting

Help the model recognize specific words it might otherwise mishear - product names, technical jargon, or unusual spellings (up to 100 terms):

python
result = client.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    keyterms=["ElevenLabs", "Scribe", "API"]
)

Language Detection

Automatic detection with optional language hint:

python
result = client.speech_to_text.convert(
    file=audio_file,
    model_id="scribe_v2",
    language_code="eng"  # ISO 639-1 or ISO 639-3 code
)

print(f"Detected: {result.language_code} ({result.language_probability:.0%})")

Supported Formats

Audio: MP3, WAV, M4A, FLAC, OGG, WebM, AAC, AIFF, Opus Video: MP4, AVI, MKV, MOV, WMV, FLV, WebM, MPEG, 3GPP

Limits: Up to 3GB file size, 10 hours duration

Response Format

json
{
  "text": "The full transcription text",
  "language_code": "eng",
  "language_probability": 0.98,
  "words": [
    { "text": "The", "start": 0.0, "end": 0.15, "type": "word", "speaker_id": "speaker_0" },
    { "text": " ", "start": 0.15, "end": 0.16, "type": "spacing", "speaker_id": "speaker_0" }
  ]
}

Word types:

  • word - An actual spoken word
  • spacing - Whitespace between words (useful for precise timing)
  • audio_event - Non-speech sounds the model detected (laughter, applause, music, etc.)

Error Handling

python
try:
    result = client.speech_to_text.convert(file=audio_file, model_id="scribe_v2")
except Exception as e:
    print(f"Transcription failed: {e}")

Common errors:

  • 401: Invalid API key
  • 422: Invalid parameters
  • 429: Rate limit exceeded

Tracking Costs

Monitor usage via request-id response header:

python
response = client.speech_to_text.convert.with_raw_response(file=audio_file, model_id="scribe_v2")
result = response.parse()
print(f"Request ID: {response.headers.get('request-id')}")
Show full SKILL.md (175 more words)Show less

Real-Time Streaming

For live transcription with ultra-low latency (~150ms), use the real-time API. The real-time API produces two types of transcripts:

  • Partial transcripts: Interim results that update frequently as audio is processed - use these for live feedback (e.g., showing text as the user speaks)
  • Committed transcripts: Final, stable results after you "commit" - use these as the source of truth for your application

A "commit" tells the model to finalize the current segment. You can commit manually (e.g., when the user pauses) or use Voice Activity Detection (VAD) to auto-commit on silence.

Python (Server-Side)
python
import asyncio
from elevenlabs import ElevenLabs

client = ElevenLabs()

async def transcribe_realtime():
    async with client.speech_to_text.realtime.connect(
        model_id="scribe_v2_realtime",
        include_timestamps=True,
    ) as connection:
        await connection.stream_url("https://example.com/audio.mp3")

        async for event in connection:
            if event.type == "partial_transcript":
                print(f"Partial: {event.text}")
            elif event.type == "committed_transcript":
                print(f"Final: {event.text}")

asyncio.run(transcribe_realtime())
JavaScript (Client-Side with React)
typescript
import { useScribe } from "@elevenlabs/react";

function TranscriptionComponent() {
  const [transcript, setTranscript] = useState("");

  const scribe = useScribe({
    modelId: "scribe_v2_realtime",
    onPartialTranscript: (data) => console.log("Partial:", data.text),
    onCommittedTranscript: (data) => setTranscript((prev) => prev + data.text),
  });

  const start = async () => {
    // Get token from your backend (never expose API key to client)
    const { token } = await fetch("/scribe-token").then((r) => r.json());

    await scribe.connect({
      token,
      microphone: { echoCancellation: true, noiseSuppression: true },
    });
  };

  return <button onClick={start}>Start Recording</button>;
}
Commit Strategies
StrategyDescription
ManualYou call commit() when ready - use for file processing or when you control the audio segments
VADVoice Activity Detection auto-commits when silence is detected - use for live microphone input
javascript
// VAD configuration
const connection = await client.speechToText.realtime.connect({
  modelId: 'scribe_v2_realtime',
  vad: {
    silenceThresholdSecs: 1.5,
    threshold: 0.4
  }
});
Event Types
EventDescription
partial_transcriptLive interim results
committed_transcriptFinal results after commit
committed_transcript_with_timestampsFinal with word timing
errorError occurred

See real-time references for complete documentation.

References

© tadaspetra, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in .agents/skills/speech-to-text of tadaspetra/loop.

  • SKILL.md
  • references/installation.md
  • references/realtime-client-side.md
  • references/realtime-commit-strategies.md
  • references/realtime-events.md
  • references/realtime-server-side.md
  • references/transcription-options.md

Open the folder on GitHubat commit 452e950

Used in 2 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in tadaspetra/loop, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Speech To Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speech To Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speech To Text this skilltadaspetra/loop2962 repos~2kAutomated safety check: PassMIT
Speech Engineelevenlabs/skills482—~2.5kAutomated safety check: WarnMIT
Local AI Useamd/skills408—~5kAutomated safety check: NotesMIT
Summarize Callreysu/ai-life-skills270—~3.8kAutomated safety check: NotesMIT
C Voicedaxaur/openpaw173—~401Automated safety check: PassMIT
Elevenlabs Core Workflow Bjeremylongshore/tons-of-skills-marketplace2.8k—~1.6kAutomated safety check: PassMIT

Similar skills

  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated 2 days ago
    Media & CreativeAuto-check: warnings
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    408 GitHub stars~5k tokensUpdated 2 days ago
    Media & CreativeAuto-check: notes
  • Summarize Call

    reysu/ai-life-skills

    Transcribe a call recording with speaker diarization, summarize it, and create Obsidian vault notes (call note, transcript, person notes for participants).

    270 GitHub stars~3.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • C Voice

    daxaur/openpaw

    Convert speech to text using sag (ElevenLabs STT) and synthesize speech using say (macOS built-in TTS).

    173 GitHub stars~401 tokensUpdated 4 mo ago
    Media & CreativeAuto-check passed
  • Elevenlabs Core Workflow B

    jeremylongshore/tons-of-skills-marketplace

    Implement ElevenLabs speech-to-speech, sound effects, audio isolation, and speech-to-text.

    2.8k GitHub stars~1.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Video Translator

    shang-zhu/violin

    Dub a video into another language and generate subtitles using the default Together + Cartesia stack.

    1.1k GitHub stars~1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes

More from tadaspetra/loop

  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Auto-check passed
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Auto-check passed
  • Agents

    tadaspetra/loop

    Build voice AI agents with ElevenLabs. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed

Questions about Speech To Text

What does Speech To Text do?

Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop. Speech To Text is an agent skill from tadaspetra/loop. Transcribe audio to text using ElevenLabs Scribe v2.

When should I use Speech To Text?

Speech To Text fits situations like: converting audio/video to text; generating subtitles; transcribing meetings; processing spoken content.

How do I install Speech To Text in Claude Code?

Run `npx skills add tadaspetra/loop --skill speech-to-text -a claude-code`. Or copy the skill folder (.agents/skills/speech-to-text in tadaspetra/loop) into .claude/skills/speech-to-text in your project. Claude Code loads it when a task matches its description.

How do I install Speech To Text in Codex?

Run `npx skills add tadaspetra/loop --skill speech-to-text -a codex`. Or copy the skill folder (.agents/skills/speech-to-text in tadaspetra/loop) into .agents/skills/speech-to-text in your project. Codex loads it when a task matches its description.

Can I use Speech To Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tadaspetra/loop --skill speech-to-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-to-text, .gemini/skills/speech-to-text, .github/skills/speech-to-text and .opencode/skills/speech-to-text in your project.

What does Speech To Text need to run?

Going by SKILL.md and its folder, Speech To Text needs the command-line tools its instructions call (curl) and credentials named ELEVENLABS_API_KEY. Our summary lists: Python 3; A credential in ELEVENLABS_API_KEY. Compatibility (from SKILL.md): Requires internet access and an ElevenLabs API key (ELEVENLABS_API_KEY)..

Does Speech To Text access the network?

SKILL.md names 1 domain. In commands or code: api.elevenlabs.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Speech To Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Speech To Text use?

Speech To Text is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Speech To Text use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.9k tokens, read only when the agent opens those files.

What are the alternatives to Speech To Text?

Skills that share tags, products or a category with Speech To Text: Speech Engine (elevenlabs/skills, 482 stars), Local AI Use (amd/skills, 408 stars), Summarize Call (reysu/ai-life-skills, 270 stars) and C Voice (daxaur/openpaw, 173 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speech To Text?

tadaspetra (a GitHub user) maintains it in tadaspetra/loop, which has 296 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 2, 2026.

Source: tadaspetra/loop on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.