Agent skill

Azure Text To Speech

by calesthio in calesthio/OpenMontage

Generate neural narration audio using Azure AI Speech (REST text-to-speech).

MITAuto-check passedMedia & Creative

Install Azure Text To Speech

skills CLI
$ npx skills add calesthio/OpenMontage --skill azure-text-to-speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage azure-text-to-speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/azure-text-to-speech .claude/skills/azure-text-to-speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azure-text-to-speech
GitHub stars
66k
Token cost
~1.4k tokens
SKILL.md length
466 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
MIT

At a glance

Generate neural narration audio using Azure AI Speech (REST text-to-speech).

  • Synthesizing voiceovers
  • SKILL.md covers Setup, Using it in a pipeline, Voice selection and Parameters that matter, plus 2 more sections
  • Needs AZURE_SPEECH_KEY
  • Narration in OpenMontage

What it does

Azure Text To Speech is an agent skill from calesthio/OpenMontage. Generate neural narration audio using Azure AI Speech (REST text-to-speech). Use when synthesizing voiceovers or narration in OpenMontage. Optional cloud TTS provider — preferred when AZURESPEECHKEY is configured; the local pipertts remains the default offline path. Shares one Speech resource with azurestt.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires internet access and an Azure AI Speech resource (AZURESPEECHKEY + AZURESPEECHREGION).

It sits in Media & Creative, covering Text to speech and voice. It works with Microsoft Azure and Azure AI Speech. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is MIT.

When your agent uses it

  • Synthesizing voiceovers
  • Narration in OpenMontage

Example prompts

  • “/azure-text-to-speech”

Requirements

  • Python 3
  • A credential in AZURE_SPEECH_KEY
  • Compatibility (from SKILL.md): Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • learn.microsoft.com
    • speech.microsoft.com
    • portal.azure.com
    • azure.microsoft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_SPEECH_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).

    From compatibility in the SKILL.md frontmatter.

Context cost

Azure Text To Speech loads about 1.4k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 466 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its MIT licence (© calesthio). 466 words, ~1,361 tokens.

Download SKILL.mdSave it as .claude/skills/azure-text-to-speech/SKILL.md (or your agent's skills folder).
name
azure-text-to-speech
description
Generate neural narration audio using Azure AI Speech (REST text-to-speech). Use when synthesizing voiceovers or narration in OpenMontage. Optional cloud TTS provider — preferred when AZURE_SPEECH_KEY is configured; the local piper_tts remains the default offline path. Shares one Speech resource with azure_stt.
compatibility
Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).
license
MIT

Azure AI Speech — Text-to-Speech

Generate narration with Azure neural TTS — high-quality multilingual voices, SSML prosody control, and express-as styles, served synchronously by the REST /cognitiveservices/v1 endpoint (no token exchange, Blob storage, or job polling). In OpenMontage this is exposed through the azure_tts tool (capability=tts, provider=azure). It is an optional cloud TTS provider — when AZURE_SPEECH_KEY is configured, prefer it for high-quality cloud narration. The local piper_tts remains the default offline path and the fallback when Azure is unavailable; elevenlabs_tts remains the choice for voice cloning.

Docs: REST text to speech · Voice gallery

Setup

Same Speech resource as azure_stt — one key/region unlocks both directions (STT and TTS). Create a Speech resource in the Azure portal; copy the key and region from its Keys and Endpoint page.

bash
export AZURE_SPEECH_KEY=your_speech_resource_key
export AZURE_SPEECH_REGION=eastus        # your resource's region
# export AZURE_TTS_ENDPOINT=https://...  # optional: full custom TTS host
#   (the TTS host is https://<region>.tts.speech.microsoft.com — a different
#    subdomain than the STT endpoint, hence the separate override var)

azure_tts reports AVAILABLE once AZURE_SPEECH_KEY plus either AZURE_SPEECH_REGION or AZURE_TTS_ENDPOINT are set.

Using it in a pipeline

Route through tts_selector as usual (it auto-discovers azure_tts), or call the provider tool directly when the user has approved Azure:

python
from tools.tool_registry import registry
registry.discover()
tts = registry._tools["azure_tts"]

result = tts.execute({
    "text": "Every design decision in this dashboard has a reason.",
    "voice": "andrew",                 # alias or full Azure short name
    "rate": "-4%",                     # slightly slower for narration
    # "style": "narration-professional",  # for voices that support styles
    "output_path": "projects/my-video/assets/audio/seg_001.mp3",
    "output_format": "mp3",            # or "wav" (48kHz PCM) for mixing
})

If azure_tts is unavailable (no key) or errors, fall back per its declared chain: elevenlabs_tts → openai_tts → piper_tts.

Voice selection

Curated shortlist (aliases accepted by the voice param):

AliasVoiceCharacter
andrewen-US-AndrewMultilingualNeuralwarm, confident, conversational — the default; founder/explainer register
brandonen-US-BrandonMultilingualNeuraldeeper, measured
avaen-US-AvaMultilingualNeuralconfident, bright female
guyen-US-GuyNeuralauthoritative
jennyen-US-JennyNeuralfriendly, clear

Any valid Azure voice short name may be passed verbatim (e.g. de-DE-KatjaNeural); the Multilingual voices handle non-English text well — set locale to match the text's language for correct SSML.

Show full SKILL.md (221 more words)Show less

Parameters that matter

  • rate / pitch — SSML prosody. Narration usually reads best slightly slowed ("-4%" to "-8%"); leave pitch at "0%" unless correcting a voice.
  • style — express-as style for voices that support it (narration-professional, calm, newscast). Unsupported styles are silently ignored by Azure, so listen to a sample before batch runs.
  • output_format — mp3 (48kHz/192kbit) for delivery, wav (48kHz PCM) when the segment feeds audio_mixer for further processing.
  • Determinism: a fixed voice + SSML re-renders effectively identical audio — safe to regenerate individual segments without re-recording the whole set.

Cost

Azure neural TTS Standard tier bills roughly $16 per 1M characters (~$0.016 per 1k chars; a 150-word narration segment ≈ $0.015). The tool reports per-call cost_usd for the cost tracker. See Azure AI Speech pricing for current rates.

Limits & tips

  • One execute call = one narration segment. Generate per script section (the asset stage convention) rather than one giant paragraph — smaller segments align cleanly to scene timings and are cheap to regenerate.
  • The synchronous endpoint caps a request at 10 minutes of audio — far above any segment OpenMontage generates.
  • Text is XML-escaped automatically; do not pre-escape or wrap in SSML — pass plain text plus the rate/pitch/style params.
  • Verify quality: listen to the first generated segment before batch-running a full script (voice/style fit is a creative decision — surface it at the proposal stage per the Decision Communication Contract).

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/azure-text-to-speech of calesthio/OpenMontage.

Open the folder on GitHubat commit 9327439

Compare with similar skills

Azure Text To Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure Text To Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure Text To Speech this skillcalesthio/OpenMontage66k—~1.4kAutomated safety check: PassMIT
Video Podcast Maker LiteAgents365-ai/video-podcast-maker1.7k—~4.5kAutomated safety check: PassMIT
Azure Speech To Text REST Pymicrosoft/skills3.1k5 repos~3kAutomated safety check: PassMIT
Azure AImicrosoft/GitHub-Copilot-for-Azure2551 repos~852Automated safety check: PassMIT
Voice Clone Ttsnpc-live/clawfirm156—~1.1kAutomated safety check: PassNone
Eardraftlawve-ai/awesome-legal-skills847—~3.3kAutomated safety check: PassMIT

Similar skills

  • Video Podcast Maker Lite

    Agents365-ai/video-podcast-maker

    Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion.

    1.7k GitHub stars~4.5k tokensUpdated 10 days ago
    Media & CreativeAuto-check passed
  • Official

    Azure Speech to Text REST API for short audio (Python). An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3k tokens
    Media & CreativeAuto-check passed
  • Azure AI

    microsoft/GitHub-Copilot-for-Azure

    Official

    A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.

    255 GitHub starsUsed in 1 repo~852 tokens
    Media & CreativeAuto-check passed
  • Voice Clone Tts

    npc-live/clawfirm

    声纹克隆和语音合成。上传音频样本克隆声纹,用克隆声纹或预设声纹生成语音。支持多个后端:MiniMax、ElevenLabs、Fish Audio、Azure TTS、OpenAI TTS。支持情绪控制、语速调整、批量生成。触发词:语音合成、TTS、声纹克隆、voice clone、text to speech、配音、旁白。

    156 GitHub stars~1.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Eardraft

    lawve-ai/awesome-legal-skills

    Transforms reading-oriented prose into listening-optimized text for flat, neutral vocal delivery — TTS, podcasts, audiobooks, CLE audio.

    847 GitHub stars~3.3k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Web Video Presentation

    ConardLi/garden-skills

    Turns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration.

    13k GitHub stars~3.5k tokensUpdated today
    Media & CreativeAuto-check passed

More from calesthio/OpenMontage

All 41 skills in this repo
  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    66k GitHub stars~841 tokensUpdated 8 days ago
    Auto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 8 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    66k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    66k GitHub stars~2k tokensUpdated 8 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    66k GitHub stars~855 tokensUpdated 8 days ago
    Auto-check: notes

Questions about Azure Text To Speech

What does Azure Text To Speech do?

Generate neural narration audio using Azure AI Speech (REST text-to-speech). Azure Text To Speech is an agent skill from calesthio/OpenMontage. Generate neural narration audio using Azure AI Speech (REST text-to-speech).

When should I use Azure Text To Speech?

Azure Text To Speech fits situations like: synthesizing voiceovers; narration in OpenMontage.

How do I install Azure Text To Speech in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill azure-text-to-speech -a claude-code`. Or copy the skill folder (.agents/skills/azure-text-to-speech in calesthio/OpenMontage) into .claude/skills/azure-text-to-speech in your project. Claude Code loads it when a task matches its description.

How do I install Azure Text To Speech in Codex?

Run `npx skills add calesthio/OpenMontage --skill azure-text-to-speech -a codex`. Or copy the skill folder (.agents/skills/azure-text-to-speech in calesthio/OpenMontage) into .agents/skills/azure-text-to-speech in your project. Codex loads it when a task matches its description.

Can I use Azure Text To Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill azure-text-to-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azure-text-to-speech, .gemini/skills/azure-text-to-speech, .github/skills/azure-text-to-speech and .opencode/skills/azure-text-to-speech in your project.

What does Azure Text To Speech need to run?

Going by SKILL.md and its folder, Azure Text To Speech needs credentials named AZURE_SPEECH_KEY. Our summary lists: Python 3; A credential in AZURE_SPEECH_KEY. Compatibility (from SKILL.md): Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION)..

Does Azure Text To Speech access the network?

SKILL.md names 4 domains. As links in the text: learn.microsoft.com, speech.microsoft.com, portal.azure.com and azure.microsoft.com. This is read from the text; nothing was executed.

Is Azure Text To Speech safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Azure Text To Speech use?

Azure Text To Speech is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure Text To Speech use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Azure Text To Speech?

Skills that share tags, products or a category with Azure Text To Speech: Video Podcast Maker Lite (Agents365-ai/video-podcast-maker, 1.7k stars), Azure Speech To Text REST Py (microsoft/skills, 3.1k stars), Azure AI (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Voice Clone Tts (npc-live/clawfirm, 156 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure Text To Speech?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.