Agent skill

Azure Speech To Text

by calesthio in calesthio/OpenMontage

Transcribe audio to text using Azure AI Speech (Fast Transcription REST API).

MITAuto-check passedMedia & Creative

Install Azure Speech To Text

skills CLI
$ npx skills add calesthio/OpenMontage --skill azure-speech-to-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage azure-speech-to-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/azure-speech-to-text .claude/skills/azure-speech-to-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azure-speech-to-text
GitHub stars
66k
Token cost
~1.3k tokens
SKILL.md length
423 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
MIT

At a glance

Transcribe audio to text using Azure AI Speech (Fast Transcription REST API).

  • Converting audio/video to text
  • SKILL.md covers Why Fast Transcription (not…, Setup, Using it in a pipeline and Parameters that matter, plus 2 more sections
  • Needs AZURE_SPEECH_KEY
  • Generating subtitles

What it does

Azure Speech To Text is an agent skill from calesthio/OpenMontage. Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content in OpenMontage. Optional cloud STT provider — preferred when AZURESPEECHKEY is configured; the local faster-whisper transcriber is the default offline path.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires internet access and an Azure AI Speech resource (AZURESPEECHKEY + AZURESPEECHREGION).

It sits in Media & Creative, covering Transcription. It works with Azure AI Speech, Whisper and Microsoft Azure. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is MIT.

When your agent uses it

  • Converting audio/video to text
  • Generating subtitles
  • Processing spoken content in OpenMontage

Example prompts

  • “/azure-speech-to-text”

Requirements

  • Python 3
  • A credential in AZURE_SPEECH_KEY
  • Compatibility (from SKILL.md): Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash, python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • learn.microsoft.com
    • portal.azure.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_SPEECH_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).

    From compatibility in the SKILL.md frontmatter.

Context cost

Azure Speech To Text loads about 1.3k tokens when it runs. Until then it costs about 88 tokens; SKILL.md has 423 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its MIT licence (© calesthio). 423 words, ~1,316 tokens.

Download SKILL.mdSave it as .claude/skills/azure-speech-to-text/SKILL.md (or your agent's skills folder).
name
azure-speech-to-text
description
Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Use when converting audio/video to text, generating subtitles, or processing spoken content in OpenMontage. Optional cloud STT provider — preferred when AZURE_SPEECH_KEY is configured; the local faster-whisper `transcriber` is the default offline path.
compatibility
Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION).
license
MIT

Azure AI Speech — Speech-to-Text

Transcribe audio to text with Azure Fast Transcription — synchronous, word-level timestamps, speaker diarization, and multi-language identification. In OpenMontage this is exposed through the azure_stt tool (capability=analysis, provider=azure). It is an optional cloud STT provider — when AZURE_SPEECH_KEY is configured, prefer it for cloud transcription. The local transcriber tool (faster-whisper) remains the default offline path and the fallback when Azure is unavailable.

Docs: Fast Transcription · Speech service overview

Why Fast Transcription (not Batch)

Azure exposes three STT surfaces. OpenMontage uses Fast Transcription because the pipeline transcribes local audio files:

SurfaceInputLatencyNeeds
Fast Transcription (used here)local file, multipart POSTsynchronous, sub-real-timekey + region
Batch Transcriptionaudio at a URL (Blob + SAS)async job + pollingBlob storage plumbing
Speech SDK (spx)mic / stream / filestreamingnative azure-cognitiveservices-speech package

Fast Transcription needs no Blob storage, no SAS URLs, and no native SDK — just requests and the two env vars.

Setup

Create a Speech resource in the Azure portal; copy the key and region from its Keys and Endpoint page.

bash
export AZURE_SPEECH_KEY=your_speech_resource_key
export AZURE_SPEECH_REGION=eastus          # your resource's region
# export AZURE_SPEECH_ENDPOINT=https://...  # optional: overrides region

azure_stt reports AVAILABLE once AZURE_SPEECH_KEY plus either AZURE_SPEECH_REGION or AZURE_SPEECH_ENDPOINT are set.

Using it in a pipeline

Prefer azure_stt over transcriber unless the run must be offline. Its output matches the transcriber schema exactly, so it is a drop-in for subtitle_gen and any stage that consumes a transcript.

python
from tools.tool_registry import registry
registry.discover()
stt = registry._tools["azure_stt"]

result = stt.execute({
    "input_path": "projects/my-video/assets/audio/narration.mp3",
    # "language": "en",          # ISO 639-1 or BCP-47 ("en-US"); omit for auto-ID
    # "diarize": True,           # speaker labels, no HuggingFace token needed
    # "max_speakers": 4,
    "output_dir": "projects/my-video/artifacts",
})
if result.success:
    segs = result.data["segments"]          # [{id,start,end,text,words:[...]}]
    words = result.data["word_timestamps"]  # flat [{word,start,end,probability}]

If azure_stt is unavailable (no key) or errors, fall back to transcriber (local whisper) — its execute signature and output are identical.

Show full SKILL.md (183 more words)Show less

Parameters that matter

  • language — pass an ISO code ("en") or a full locale ("en-US"). Pin it when you know the language; it is faster and more accurate than auto-ID.
  • candidate_locales — when language is omitted, Azure runs language identification across this shortlist. Narrow it to the languages you actually expect; a huge list slows detection and invites misclassification.
  • diarize / max_speakers — enable for multi-speaker audio (interviews, podcasts). Set max_speakers to the real upper bound.
  • profanity_filter — None | Masked (default) | Removed | Tags.

Response shape (mapped to the transcriber schema)

The raw Azure response (phrases[] with offsetMilliseconds / words[]) is converted to seconds and the OpenMontage transcript schema:

json
{
  "segments": [
    {"id": 0, "start": 0.0, "end": 2.4, "text": "Hello world",
     "speaker": 1,
     "words": [{"word": "Hello", "start": 0.0, "end": 0.5, "probability": 0.98}]}
  ],
  "word_timestamps": [{"word": "Hello", "start": 0.0, "end": 0.5, "probability": 0.98}],
  "language": "en-US",
  "duration_seconds": 2.4,
  "provider": "azure"
}

Note: Fast Transcription has no per-word confidence, so each word carries the phrase confidence in probability.

Limits & tips

  • Single file up to ~2 hours / a few hundred MB per request. For longer or bulk jobs, use Azure Batch Transcription instead.
  • Send clean audio (16 kHz+ mono is plenty). Transcode video to audio first if you only need speech — smaller upload, same result.
  • Verify timing: word timestamps drive subtitle cues in subtitle_gen. Spot-check the first and last cues against the source audio.

© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/azure-speech-to-text of calesthio/OpenMontage.

Open the folder on GitHubat commit 9327439

Compare with similar skills

Azure Speech To Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure Speech To Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure Speech To Text this skillcalesthio/OpenMontage66k—~1.3kAutomated safety check: PassMIT
Azure Speech To Text REST Pymicrosoft/skills3.1k5 repos~3kAutomated safety check: PassMIT
Deepgram Migration Deep Divejeremylongshore/tons-of-skills-marketplace2.8k—~3.3kAutomated safety check: PassMIT
Azure AI Voicelive Pymicrosoft/skills3.1k—~2.9kAutomated safety check: PassMIT
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
ShortsAgriciDaniel/claude-shorts219—~3.2kAutomated safety check: NotesMIT

Similar skills

  • Official

    Azure Speech to Text REST API for short audio (Python). An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3k tokens
    Media & CreativeAuto-check passed
  • Deepgram Migration Deep Dive

    jeremylongshore/tons-of-skills-marketplace

    Deep dive into migrating to Deepgram from other transcription providers.

    2.8k GitHub stars~3.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Azure AI Voicelive Py

    microsoft/skills

    Official

    Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).

    3.1k GitHub stars~2.9k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Shorts

    AgriciDaniel/claude-shorts

    Interactive longform-to-shortform video creator. An agent skill from AgriciDaniel/claude-shorts.

    219 GitHub stars~3.2k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Video To Subtitle Summary

    imlewc/video-to-subtitle-summary-skill

    A skill your agent uses when user provides a short video platform URL or local video/audio file and wants subtitles/AI summary, or when user asks to list their own AI Douyin historical tasks.

    218 GitHub stars~4.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes

More from calesthio/OpenMontage

All 41 skills in this repo
  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    66k GitHub stars~841 tokensUpdated 8 days ago
    Auto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 8 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    66k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    66k GitHub stars~2k tokensUpdated 8 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    66k GitHub stars~855 tokensUpdated 8 days ago
    Auto-check: notes

Questions about Azure Speech To Text

What does Azure Speech To Text do?

Transcribe audio to text using Azure AI Speech (Fast Transcription REST API). Azure Speech To Text is an agent skill from calesthio/OpenMontage. Transcribe audio to text using Azure AI Speech (Fast Transcription REST API).

When should I use Azure Speech To Text?

Azure Speech To Text fits situations like: converting audio/video to text; generating subtitles; processing spoken content in OpenMontage.

How do I install Azure Speech To Text in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill azure-speech-to-text -a claude-code`. Or copy the skill folder (.agents/skills/azure-speech-to-text in calesthio/OpenMontage) into .claude/skills/azure-speech-to-text in your project. Claude Code loads it when a task matches its description.

How do I install Azure Speech To Text in Codex?

Run `npx skills add calesthio/OpenMontage --skill azure-speech-to-text -a codex`. Or copy the skill folder (.agents/skills/azure-speech-to-text in calesthio/OpenMontage) into .agents/skills/azure-speech-to-text in your project. Codex loads it when a task matches its description.

Can I use Azure Speech To Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill azure-speech-to-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azure-speech-to-text, .gemini/skills/azure-speech-to-text, .github/skills/azure-speech-to-text and .opencode/skills/azure-speech-to-text in your project.

What does Azure Speech To Text need to run?

Going by SKILL.md and its folder, Azure Speech To Text needs credentials named AZURE_SPEECH_KEY. Our summary lists: Python 3; A credential in AZURE_SPEECH_KEY. Compatibility (from SKILL.md): Requires internet access and an Azure AI Speech resource (AZURE_SPEECH_KEY + AZURE_SPEECH_REGION)..

Does Azure Speech To Text access the network?

SKILL.md names 2 domains. As links in the text: learn.microsoft.com and portal.azure.com. This is read from the text; nothing was executed.

Is Azure Speech To Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Azure Speech To Text use?

Azure Speech To Text is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure Speech To Text use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Azure Speech To Text?

Skills that share tags, products or a category with Azure Speech To Text: Azure Speech To Text REST Py (microsoft/skills, 3.1k stars), Deepgram Migration Deep Dive (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Azure AI Voicelive Py (microsoft/skills, 3.1k stars) and 9Router Speech-to-Text (decolua/9router, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure Speech To Text?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.