Hyperframes Media
chmonitor/chmonitor
Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…
Generate speech audio from text using HeyGen's Starfish TTS model.
$ npx skills add calesthio/OpenMontage --skill text-to-speech -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calesthio/OpenMontage text-to-speech --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/text-to-speech .claude/skills/text-to-speech && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "text-to-speech" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/text-to-speech into .claude/skills/text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-to-speech", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/text-to-speechType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calesthio/OpenMontage --skill text-to-speech -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calesthio/OpenMontage text-to-speech --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/text-to-speech .agents/skills/text-to-speech && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "text-to-speech" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/text-to-speech into .agents/skills/text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-to-speech", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/OpenMontage --skill text-to-speech -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calesthio/OpenMontage text-to-speech --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/text-to-speech .cursor/skills/text-to-speech && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "text-to-speech" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/text-to-speech into .cursor/skills/text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-to-speech", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calesthio/OpenMontage.git --path .agents/skills/text-to-speech--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calesthio/OpenMontage --skill text-to-speech -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calesthio/OpenMontage text-to-speech --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/text-to-speech .gemini/skills/text-to-speech && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "text-to-speech" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/text-to-speech into .gemini/skills/text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-to-speech", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calesthio/OpenMontage text-to-speechInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calesthio/OpenMontage --skill text-to-speech -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/text-to-speech .github/skills/text-to-speech && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "text-to-speech" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/text-to-speech into .github/skills/text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-to-speech", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/OpenMontage --skill text-to-speech -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calesthio/OpenMontage text-to-speech --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/text-to-speech .opencode/skills/text-to-speech && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "text-to-speech" agent skill from https://github.com/calesthio/OpenMontage/tree/main/.agents/skills/text-to-speech into .opencode/skills/text-to-speech/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "text-to-speech", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
text-to-speechGenerate speech audio from text using HeyGen's Starfish TTS model.
Text To Speech is an agent skill from calesthio/OpenMontage. Generate speech audio from text using HeyGen's Starfish TTS model. Use when: (1) Generating standalone speech audio files from text, (2) Converting text to speech with voice selection, speed, and pitch control, (3) Creating audio for voiceovers, narration, or podcasts, (4) Working with HeyGen's /v1/audio endpoints, (5) Listing available TTS voices by language or gender.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Text to speech and voice. It works with HeyGen, ElevenLabs and Model Context Protocol. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is AGPL-3.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
mcp__heygen__*From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.heygen.comresource.heygen.airesource2.heygen.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HEYGEN_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Text To Speech loads about 2.4k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 489 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from calesthio/OpenMontage at commit 9327439, republished under its AGPL-3.0 licence (© calesthio). 489 words, ~2,442 tokens.
.claude/skills/text-to-speech/SKILL.md (or your agent's skills folder).Generate speech audio files from text using HeyGen's in-house Starfish TTS model. This skill is for standalone audio generation — separate from video creation.
All requests require the X-Api-Key header. Set the HEYGEN_API_KEY environment variable.
curl -X GET "https://api.heygen.com/v1/audio/voices" \
-H "X-Api-Key: $HEYGEN_API_KEY"If HeyGen MCP tools are available (mcp__heygen__*), prefer them over direct HTTP API calls.
| Task | MCP Tool | Fallback (Direct API) |
|---|---|---|
| List TTS voices | mcp__heygen__list_audio_voices | GET /v1/audio/voices |
| Generate speech audio | mcp__heygen__text_to_speech | POST /v1/audio/text_to_speech |
mcp__heygen__list_audio_voices (or GET /v1/audio/voices)mcp__heygen__text_to_speech (or POST /v1/audio/text_to_speech) with text and voice_idaudio_url to download or play the audioRetrieve voices compatible with the Starfish TTS model.
Note: This uses
GET /v1/audio/voices— a different endpoint from the video voices API (GET /v2/voices). Not all video voices support Starfish TTS.
curl -X GET "https://api.heygen.com/v1/audio/voices" \
-H "X-Api-Key: $HEYGEN_API_KEY"interface TTSVoice {
voice_id: string;
language: string;
gender: "female" | "male" | "unknown";
name: string;
preview_audio_url: string | null;
support_pause: boolean;
support_locale: boolean;
type: string;
}
interface TTSVoicesResponse {
error: null | string;
data: {
voices: TTSVoice[];
};
}
async function listTTSVoices(): Promise<TTSVoice[]> {
const response = await fetch("https://api.heygen.com/v1/audio/voices", {
headers: { "X-Api-Key": process.env.HEYGEN_API_KEY! },
});
const json: TTSVoicesResponse = await response.json();
if (json.error) {
throw new Error(json.error);
}
return json.data.voices;
}import requests
import os
def list_tts_voices() -> list:
response = requests.get(
"https://api.heygen.com/v1/audio/voices",
headers={"X-Api-Key": os.environ["HEYGEN_API_KEY"]}
)
data = response.json()
if data.get("error"):
raise Exception(data["error"])
return data["data"]["voices"]{
"error": null,
"data": {
"voices": [
{
"voice_id": "f38a635bee7a4d1f9b0a654a31d050d2",
"name": "Chill Brian",
"language": "English",
"gender": "male",
"preview_audio_url": "https://resource.heygen.ai/text_to_speech/WpSDQvmLGXEqXZVZQiVeg6.mp3",
"support_pause": true,
"support_locale": false,
"type": "public"
}
]
}
}Convert text to speech audio using a specified voice.
POST https://api.heygen.com/v1/audio/text_to_speech
| Field | Type | Req | Description |
|---|---|---|---|
text | string | Y | Text content to convert to speech |
voice_id | string | Y | Voice ID from GET /v1/audio/voices |
speed | number | Speech speed, 0.5-1.5 (default: 1) | |
pitch | integer | Voice pitch, -50 to 50 (default: 0) | |
locale | string | Accent/locale for multilingual voices (e.g., en-US, pt-BR) | |
elevenlabs_settings | object | Advanced settings for ElevenLabs voices |
| Field | Type | Description |
|---|---|---|
model | string | Model selection (eleven_v3, eleven_turbo_v2_5, etc.) |
similarity_boost | number | Voice similarity, 0.0-1.0 |
stability | number | Output consistency, 0.0-1.0 |
style | number | Style intensity, 0.0-1.0 |
curl -X POST "https://api.heygen.com/v1/audio/text_to_speech" \
-H "X-Api-Key: $HEYGEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"text": "Hello! Welcome to our product demo.",
"voice_id": "YOUR_VOICE_ID",
"speed": 1.0
}'interface TTSRequest {
text: string;
voice_id: string;
speed?: number;
pitch?: number;
locale?: string;
elevenlabs_settings?: {
model?: string;
similarity_boost?: number;
stability?: number;
style?: number;
};
}
interface WordTimestamp {
word: string;
start: number;
end: number;
}
interface TTSResponse {
error: null | string;
data: {
audio_url: string;
duration: number;
request_id: string;
word_timestamps: WordTimestamp[];
};
}
async function textToSpeech(request: TTSRequest): Promise<TTSResponse["data"]> {
const response = await fetch(
"https://api.heygen.com/v1/audio/text_to_speech",
{
method: "POST",
headers: {
"X-Api-Key": process.env.HEYGEN_API_KEY!,
"Content-Type": "application/json",
},
body: JSON.stringify(request),
}
);
const json: TTSResponse = await response.json();
if (json.error) {
throw new Error(json.error);
}
return json.data;
}import requests
import os
def text_to_speech(
text: str,
voice_id: str,
speed: float = 1.0,
pitch: int = 0,
locale: str | None = None,
) -> dict:
payload = {
"text": text,
"voice_id": voice_id,
"speed": speed,
"pitch": pitch,
}
if locale:
payload["locale"] = locale
response = requests.post(
"https://api.heygen.com/v1/audio/text_to_speech",
headers={
"X-Api-Key": os.environ["HEYGEN_API_KEY"],
"Content-Type": "application/json",
},
json=payload,
)
data = response.json()
if data.get("error"):
raise Exception(data["error"])
return data["data"]{
"error": null,
"data": {
"audio_url": "https://resource2.heygen.ai/text_to_speech/.../id=365d46bb.wav",
"duration": 5.526,
"request_id": "p38QJ52hfgNlsYKZZmd9",
"word_timestamps": [
{ "word": "<start>", "start": 0.0, "end": 0.0 },
{ "word": "Hey", "start": 0.079, "end": 0.219 },
{ "word": "there,", "start": 0.239, "end": 0.459 },
{ "word": "<end>", "start": 5.526, "end": 5.526 }
]
}
}const result = await textToSpeech({
text: "Welcome to our quarterly earnings call.",
voice_id: "YOUR_VOICE_ID",
});
console.log(`Audio URL: ${result.audio_url}`);
console.log(`Duration: ${result.duration}s`);const result = await textToSpeech({
text: "We're thrilled to announce our newest feature!",
voice_id: "YOUR_VOICE_ID",
speed: 1.1,
});const result = await textToSpeech({
text: "Bem-vindo ao nosso produto.",
voice_id: "MULTILINGUAL_VOICE_ID",
locale: "pt-BR",
});async function generateSpeech(text: string, language: string): Promise<string> {
const voices = await listTTSVoices();
const voice = voices.find(
(v) => v.language.toLowerCase().includes(language.toLowerCase())
);
if (!voice) {
throw new Error(`No TTS voice found for language: ${language}`);
}
const result = await textToSpeech({
text,
voice_id: voice.voice_id,
});
return result.audio_url;
}
const audioUrl = await generateSpeech("Hello and welcome!", "english");Use SSML-style break tags in your text for pauses:
word <break time="1s"/> wordRules:
s suffix: <break time="1.5s"/>For narration, create a short voice-performance plan before generating audio:
Use concrete cues, not generic instructions. "Warm but decisive; pause before the contrast; slow down on the final sentence" is useful. "Sound natural" is not.
When the selected voice supports pauses, put the most important pauses directly in the text with break tags. Generate a sample from the most performance-heavy section first, and do not batch-generate the rest if the sample sounds flat, rushed, or ignores the intended breaks.
GET /v1/audio/voices to find compatible voices — not all voices from GET /v2/voices support Starfish TTSsupport_locale before setting a locale — only multilingual voices support locale selectionpreview_audio_url before generating (may be null for some voices)word_timestamps in the response for caption syncing or timed text overlaysword <break time="1s"/> word© calesthio, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/text-to-speech of calesthio/OpenMontage.
Open the folder on GitHubat commit 9327439
Text To Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Text To Speech this skillcalesthio/OpenMontage | 66k | — | ~2.4k | Automated safety check: Pass | AGPL-3.0 | |
| Hyperframes Mediachmonitor/chmonitor | 300 | 1 repos | ~2.8k | Automated safety check: Notes | GPL-3.0 | |
| Super Video MakerBomx/super-video-maker-skill | 310 | — | ~11k | Automated safety check: Notes | None | |
| Motion Videobestagentkits/motion-video-skill | 118 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Remotion ProductionDojoCodingLabs/remotion-superpowers | 132 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Story Narratorhassancs91/claude-image-generation | 102 | — | ~2.5k | Automated safety check: Pass | MIT |
chmonitor/chmonitor
Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…
Bomx/super-video-maker-skill
End-to-end AI video production skill for agentic frameworks.
bestagentkits/motion-video-skill
Produce beat-synced 1080p motion-graphic videos in HyperFrames (HTML + GSAP) with an AI voice-over, Vietnamese karaoke captions, SFX and generated music, in one of two proven styles (glass keynote…
DojoCodingLabs/remotion-superpowers
Full video production workflow for Remotion projects. An agent skill from DojoCodingLabs/remotion-superpowers.
hassancs91/claude-image-generation
Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech).
social-media-skills/skills
The AI narration / voiceover mini-skill (ElevenLabs-led). An agent skill from social-media-skills/skills.
calesthio/OpenMontage
Understand video content locally using ffmpeg frame extraction and Whisper transcription.
calesthio/OpenMontage
Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.
calesthio/OpenMontage
Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.
calesthio/OpenMontage
Create videos from a text prompt using HeyGen's Video Agent.
calesthio/OpenMontage
Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.
calesthio/OpenMontage
Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.
Works with
Categories
Generate speech audio from text using HeyGen's Starfish TTS model. Text To Speech is an agent skill from calesthio/OpenMontage. Generate speech audio from text using HeyGen's Starfish TTS model.
Text To Speech fits situations like: generating standalone speech audio files from text; converting text to speech with voice selection; creating audio for voiceovers; working with HeyGens /v1/audio endpoints.
Run `npx skills add calesthio/OpenMontage --skill text-to-speech -a claude-code`. Or copy the skill folder (.agents/skills/text-to-speech in calesthio/OpenMontage) into .claude/skills/text-to-speech in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calesthio/OpenMontage --skill text-to-speech -a codex`. Or copy the skill folder (.agents/skills/text-to-speech in calesthio/OpenMontage) into .agents/skills/text-to-speech in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill text-to-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/text-to-speech, .gemini/skills/text-to-speech, .github/skills/text-to-speech and .opencode/skills/text-to-speech in your project.
Going by SKILL.md and its folder, Text To Speech needs the command-line tools its instructions call (curl) and credentials named HEYGEN_API_KEY. Our summary lists: Python 3; A credential in HEYGEN_API_KEY. Its frontmatter pre-approves these tools: mcp__heygen__*.
SKILL.md names 3 domains. In commands or code: api.heygen.com, resource.heygen.ai and resource2.heygen.ai; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Text To Speech is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Text To Speech: Hyperframes Media (chmonitor/chmonitor, 300 stars), Super Video Maker (Bomx/super-video-maker-skill, 310 stars), Motion Video (bestagentkits/motion-video-skill, 118 stars) and Remotion Production (DojoCodingLabs/remotion-superpowers, 132 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.
Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.