Agent skill

Text To Speech

by calesthio in calesthio/OpenMontage

Generate speech audio from text using HeyGen's Starfish TTS model.

AGPL-3.0Auto-check passedMedia & Creative

Install Text To Speech

skills CLI
$ npx skills add calesthio/OpenMontage --skill text-to-speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage text-to-speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/text-to-speech .claude/skills/text-to-speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
text-to-speech
GitHub stars
66k
Token cost
~2.4k tokens
SKILL.md length
489 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Generate speech audio from text using HeyGen's Starfish TTS model.

  • Works in 4 steps: List voices with… → Pick a voice matching desired language,… → Call mcpheygentext_to_speech (or POST… → …
  • Generating standalone speech audio files from text
  • SKILL.md covers Authentication, Tool Selection, Default Workflow and List TTS Voices, plus 5 more sections
  • Calls curl; reaches api.heygen.com and resource.heygen.ai; needs HEYGEN_API_KEY

What it does

Text To Speech is an agent skill from calesthio/OpenMontage. Generate speech audio from text using HeyGen's Starfish TTS model. Use when: (1) Generating standalone speech audio files from text, (2) Converting text to speech with voice selection, speed, and pitch control, (3) Creating audio for voiceovers, narration, or podcasts, (4) Working with HeyGen's /v1/audio endpoints, (5) Listing available TTS voices by language or gender.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with HeyGen, ElevenLabs and Model Context Protocol. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is AGPL-3.0.

When your agent uses it

  • Generating standalone speech audio files from text
  • Converting text to speech with voice selection
  • Creating audio for voiceovers
  • Working with HeyGens /v1/audio endpoints

Example prompts

  • “/text-to-speech”

Requirements

  • Python 3
  • A credential in HEYGEN_API_KEY
  • Pre-approved tools (allowed-tools): mcp__heygen__*

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. List voices with mcpheygenlist_audio_voices (or GET /v1/audio/voices)
  2. Pick a voice matching desired language, gender, and features
  3. Call mcpheygentext_to_speech (or POST /v1/audio/text_to_speech) with text and voice_id
  4. Use the returned audio_url to download or play the audio

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • mcp__heygen__*

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.heygen.com
    • resource.heygen.ai
    • resource2.heygen.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HEYGEN_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Text To Speech loads about 2.4k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 489 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its AGPL-3.0 licence (© calesthio). 489 words, ~2,442 tokens.

Download SKILL.mdSave it as .claude/skills/text-to-speech/SKILL.md (or your agent's skills folder).
name
text-to-speech
description
Generate speech audio from text using HeyGen's Starfish TTS model. Use when: (1) Generating standalone speech audio files from text, (2) Converting text to speech with voice selection, speed, and pitch control, (3) Creating audio for voiceovers, narration, or podcasts, (4) Working with HeyGen's /v1/audio endpoints, (5) Listing available TTS voices by language or gender.
allowed-tools
mcp__heygen__*

Text-to-Speech (HeyGen Starfish)

Generate speech audio files from text using HeyGen's in-house Starfish TTS model. This skill is for standalone audio generation — separate from video creation.

Authentication

All requests require the X-Api-Key header. Set the HEYGEN_API_KEY environment variable.

bash
curl -X GET "https://api.heygen.com/v1/audio/voices" \
  -H "X-Api-Key: $HEYGEN_API_KEY"

Tool Selection

If HeyGen MCP tools are available (mcp__heygen__*), prefer them over direct HTTP API calls.

TaskMCP ToolFallback (Direct API)
List TTS voicesmcp__heygen__list_audio_voicesGET /v1/audio/voices
Generate speech audiomcp__heygen__text_to_speechPOST /v1/audio/text_to_speech

Default Workflow

  1. List voices with mcp__heygen__list_audio_voices (or GET /v1/audio/voices)
  2. Pick a voice matching desired language, gender, and features
  3. Call mcp__heygen__text_to_speech (or POST /v1/audio/text_to_speech) with text and voice_id
  4. Use the returned audio_url to download or play the audio

List TTS Voices

Retrieve voices compatible with the Starfish TTS model.

Note: This uses GET /v1/audio/voices — a different endpoint from the video voices API (GET /v2/voices). Not all video voices support Starfish TTS.

curl
bash
curl -X GET "https://api.heygen.com/v1/audio/voices" \
  -H "X-Api-Key: $HEYGEN_API_KEY"
TypeScript
typescript
interface TTSVoice {
  voice_id: string;
  language: string;
  gender: "female" | "male" | "unknown";
  name: string;
  preview_audio_url: string | null;
  support_pause: boolean;
  support_locale: boolean;
  type: string;
}

interface TTSVoicesResponse {
  error: null | string;
  data: {
    voices: TTSVoice[];
  };
}

async function listTTSVoices(): Promise<TTSVoice[]> {
  const response = await fetch("https://api.heygen.com/v1/audio/voices", {
    headers: { "X-Api-Key": process.env.HEYGEN_API_KEY! },
  });

  const json: TTSVoicesResponse = await response.json();

  if (json.error) {
    throw new Error(json.error);
  }

  return json.data.voices;
}
Python
python
import requests
import os

def list_tts_voices() -> list:
    response = requests.get(
        "https://api.heygen.com/v1/audio/voices",
        headers={"X-Api-Key": os.environ["HEYGEN_API_KEY"]}
    )

    data = response.json()
    if data.get("error"):
        raise Exception(data["error"])

    return data["data"]["voices"]
Response Format
json
{
  "error": null,
  "data": {
    "voices": [
      {
        "voice_id": "f38a635bee7a4d1f9b0a654a31d050d2",
        "name": "Chill Brian",
        "language": "English",
        "gender": "male",
        "preview_audio_url": "https://resource.heygen.ai/text_to_speech/WpSDQvmLGXEqXZVZQiVeg6.mp3",
        "support_pause": true,
        "support_locale": false,
        "type": "public"
      }
    ]
  }
}

Generate Speech Audio

Convert text to speech audio using a specified voice.

Endpoint

POST https://api.heygen.com/v1/audio/text_to_speech

Request Fields
FieldTypeReqDescription
textstringYText content to convert to speech
voice_idstringYVoice ID from GET /v1/audio/voices
speednumberSpeech speed, 0.5-1.5 (default: 1)
pitchintegerVoice pitch, -50 to 50 (default: 0)
localestringAccent/locale for multilingual voices (e.g., en-US, pt-BR)
elevenlabs_settingsobjectAdvanced settings for ElevenLabs voices
ElevenLabs Settings (optional)
FieldTypeDescription
modelstringModel selection (eleven_v3, eleven_turbo_v2_5, etc.)
similarity_boostnumberVoice similarity, 0.0-1.0
stabilitynumberOutput consistency, 0.0-1.0
stylenumberStyle intensity, 0.0-1.0
curl
bash
curl -X POST "https://api.heygen.com/v1/audio/text_to_speech" \
  -H "X-Api-Key: $HEYGEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Hello! Welcome to our product demo.",
    "voice_id": "YOUR_VOICE_ID",
    "speed": 1.0
  }'
TypeScript
typescript
interface TTSRequest {
  text: string;
  voice_id: string;
  speed?: number;
  pitch?: number;
  locale?: string;
  elevenlabs_settings?: {
    model?: string;
    similarity_boost?: number;
    stability?: number;
    style?: number;
  };
}

interface WordTimestamp {
  word: string;
  start: number;
  end: number;
}

interface TTSResponse {
  error: null | string;
  data: {
    audio_url: string;
    duration: number;
    request_id: string;
    word_timestamps: WordTimestamp[];
  };
}

async function textToSpeech(request: TTSRequest): Promise<TTSResponse["data"]> {
  const response = await fetch(
    "https://api.heygen.com/v1/audio/text_to_speech",
    {
      method: "POST",
      headers: {
        "X-Api-Key": process.env.HEYGEN_API_KEY!,
        "Content-Type": "application/json",
      },
      body: JSON.stringify(request),
    }
  );

  const json: TTSResponse = await response.json();

  if (json.error) {
    throw new Error(json.error);
  }

  return json.data;
}
Python
python
import requests
import os

def text_to_speech(
    text: str,
    voice_id: str,
    speed: float = 1.0,
    pitch: int = 0,
    locale: str | None = None,
) -> dict:
    payload = {
        "text": text,
        "voice_id": voice_id,
        "speed": speed,
        "pitch": pitch,
    }

    if locale:
        payload["locale"] = locale

    response = requests.post(
        "https://api.heygen.com/v1/audio/text_to_speech",
        headers={
            "X-Api-Key": os.environ["HEYGEN_API_KEY"],
            "Content-Type": "application/json",
        },
        json=payload,
    )

    data = response.json()
    if data.get("error"):
        raise Exception(data["error"])

    return data["data"]
Response Format
json
{
  "error": null,
  "data": {
    "audio_url": "https://resource2.heygen.ai/text_to_speech/.../id=365d46bb.wav",
    "duration": 5.526,
    "request_id": "p38QJ52hfgNlsYKZZmd9",
    "word_timestamps": [
      { "word": "<start>", "start": 0.0, "end": 0.0 },
      { "word": "Hey", "start": 0.079, "end": 0.219 },
      { "word": "there,", "start": 0.239, "end": 0.459 },
      { "word": "<end>", "start": 5.526, "end": 5.526 }
    ]
  }
}

Usage Examples

Basic TTS
typescript
const result = await textToSpeech({
  text: "Welcome to our quarterly earnings call.",
  voice_id: "YOUR_VOICE_ID",
});

console.log(`Audio URL: ${result.audio_url}`);
console.log(`Duration: ${result.duration}s`);
With Speed Adjustment
typescript
const result = await textToSpeech({
  text: "We're thrilled to announce our newest feature!",
  voice_id: "YOUR_VOICE_ID",
  speed: 1.1,
});
With Locale for Multilingual Voices
typescript
const result = await textToSpeech({
  text: "Bem-vindo ao nosso produto.",
  voice_id: "MULTILINGUAL_VOICE_ID",
  locale: "pt-BR",
});
Find a Voice and Generate Audio
typescript
async function generateSpeech(text: string, language: string): Promise<string> {
  const voices = await listTTSVoices();
  const voice = voices.find(
    (v) => v.language.toLowerCase().includes(language.toLowerCase())
  );

  if (!voice) {
    throw new Error(`No TTS voice found for language: ${language}`);
  }

  const result = await textToSpeech({
    text,
    voice_id: voice.voice_id,
  });

  return result.audio_url;
}

const audioUrl = await generateSpeech("Hello and welcome!", "english");

Pauses with Break Tags

Use SSML-style break tags in your text for pauses:

word <break time="1s"/> word

Rules:

  • Use seconds with s suffix: <break time="1.5s"/>
  • Must have spaces before and after the tag
  • Self-closing tag format
Show full SKILL.md (185 more words)Show less

Expressive Voice Direction

For narration, create a short voice-performance plan before generating audio:

  • narrator persona and emotional intent
  • pacing profile
  • energy curve across the script
  • where pauses should land
  • words or phrases that need emphasis

Use concrete cues, not generic instructions. "Warm but decisive; pause before the contrast; slow down on the final sentence" is useful. "Sound natural" is not.

When the selected voice supports pauses, put the most important pauses directly in the text with break tags. Generate a sample from the most performance-heavy section first, and do not batch-generate the rest if the sample sounds flat, rushed, or ignores the intended breaks.

Best Practices

  1. Use GET /v1/audio/voices to find compatible voices — not all voices from GET /v2/voices support Starfish TTS
  2. Check support_locale before setting a locale — only multilingual voices support locale selection
  3. Keep speed between 0.8-1.2 for natural-sounding output
  4. Preview voices using the preview_audio_url before generating (may be null for some voices)
  5. Use word_timestamps in the response for caption syncing or timed text overlays
  6. Use SSML break tags in your text for pauses: word <break time="1s"/> word

© calesthio, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/text-to-speech of calesthio/OpenMontage.

Open the folder on GitHubat commit 9327439

Compare with similar skills

Text To Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Text To Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Text To Speech this skillcalesthio/OpenMontage66k—~2.4kAutomated safety check: PassAGPL-3.0
Hyperframes Mediachmonitor/chmonitor3001 repos~2.8kAutomated safety check: NotesGPL-3.0
Super Video MakerBomx/super-video-maker-skill310—~11kAutomated safety check: NotesNone
Motion Videobestagentkits/motion-video-skill118—~1.5kAutomated safety check: PassMIT
Remotion ProductionDojoCodingLabs/remotion-superpowers132—~1.1kAutomated safety check: PassMIT
Story Narratorhassancs91/claude-image-generation102—~2.5kAutomated safety check: PassMIT

Similar skills

  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    300 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Motion Video

    bestagentkits/motion-video-skill

    Produce beat-synced 1080p motion-graphic videos in HyperFrames (HTML + GSAP) with an AI voice-over, Vietnamese karaoke captions, SFX and generated music, in one of two proven styles (glass keynote…

    118 GitHub stars~1.5k tokensUpdated 12 days ago
    Media & CreativeAuto-check passed
  • Remotion Production

    DojoCodingLabs/remotion-superpowers

    Full video production workflow for Remotion projects. An agent skill from DojoCodingLabs/remotion-superpowers.

    132 GitHub stars~1.1k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Story Narrator

    hassancs91/claude-image-generation

    Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech).

    102 GitHub stars~2.5k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • AI Voiceover

    social-media-skills/skills

    The AI narration / voiceover mini-skill (ElevenLabs-led). An agent skill from social-media-skills/skills.

    134 GitHub stars~1.1k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed

More from calesthio/OpenMontage

All 41 skills in this repo
  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    66k GitHub stars~841 tokensUpdated 7 days ago
    Auto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    66k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    66k GitHub stars~2k tokensUpdated 7 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    66k GitHub stars~855 tokensUpdated 7 days ago
    Auto-check: notes

Questions about Text To Speech

What does Text To Speech do?

Generate speech audio from text using HeyGen's Starfish TTS model. Text To Speech is an agent skill from calesthio/OpenMontage. Generate speech audio from text using HeyGen's Starfish TTS model.

When should I use Text To Speech?

Text To Speech fits situations like: generating standalone speech audio files from text; converting text to speech with voice selection; creating audio for voiceovers; working with HeyGens /v1/audio endpoints.

How do I install Text To Speech in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill text-to-speech -a claude-code`. Or copy the skill folder (.agents/skills/text-to-speech in calesthio/OpenMontage) into .claude/skills/text-to-speech in your project. Claude Code loads it when a task matches its description.

How do I install Text To Speech in Codex?

Run `npx skills add calesthio/OpenMontage --skill text-to-speech -a codex`. Or copy the skill folder (.agents/skills/text-to-speech in calesthio/OpenMontage) into .agents/skills/text-to-speech in your project. Codex loads it when a task matches its description.

Can I use Text To Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill text-to-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/text-to-speech, .gemini/skills/text-to-speech, .github/skills/text-to-speech and .opencode/skills/text-to-speech in your project.

What does Text To Speech need to run?

Going by SKILL.md and its folder, Text To Speech needs the command-line tools its instructions call (curl) and credentials named HEYGEN_API_KEY. Our summary lists: Python 3; A credential in HEYGEN_API_KEY. Its frontmatter pre-approves these tools: mcp__heygen__*.

Does Text To Speech access the network?

SKILL.md names 3 domains. In commands or code: api.heygen.com, resource.heygen.ai and resource2.heygen.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Text To Speech safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Text To Speech use?

Text To Speech is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Text To Speech use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Text To Speech?

Skills that share tags, products or a category with Text To Speech: Hyperframes Media (chmonitor/chmonitor, 300 stars), Super Video Maker (Bomx/super-video-maker-skill, 310 stars), Motion Video (bestagentkits/motion-video-skill, 118 stars) and Remotion Production (DojoCodingLabs/remotion-superpowers, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Text To Speech?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.