Official agent skill

Azure Speech To Text REST Py

by microsoft in microsoft/skills

Azure Speech to Text REST API for short audio (Python). An agent skill from microsoft/skills.

OfficialMITAuto-check passedMedia & Creative

Install Azure Speech To Text REST Py

skills CLI
$ npx skills add microsoft/skills --skill azure-speech-to-text-rest-py -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/skills azure-speech-to-text-rest-py --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/plugins/azure-sdk-python/skills/azure-speech-to-text-rest-py .claude/skills/azure-speech-to-text-rest-py && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azure-speech-to-text-rest-py
GitHub stars
3.1k
Used in
5 other repos
Token cost
~3k tokens
SKILL.md length
489 words
Files
2 (incl. references)
Skills in repo
150
Repo updated
First seen
Licence
MIT

At a glance

Azure Speech to Text REST API for short audio (Python). An agent skill from microsoft/skills.

  • Works in 3 steps: Azure subscription - Create one free → Speech resource - Create in Azure Portal → Get credentials - After deployment, go…
  • Simple speech recognition of audio files up to 60 seconds without the Speech SDK
  • SKILL.md covers Prerequisites, Environment Variables, Installation and Authentication & Lifecycle, plus 15 more sections
  • Calls pip; needs AZURE_SPEECH_KEY

What it does

Azure Speech To Text REST Py is an agent skill from microsoft/skills, published by the product's own GitHub organization. Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK. Triggers: "speech to text REST", "short audio transcription", "speech recognition REST API", "STT REST", "recognize speech REST". DO NOT USE FOR: Long audio (60 seconds), real-time streaming, batch transcription, custom speech models, speech translation. Use Speech SDK or Batch Transcription API instead.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/pronunciation-assessment.md`).

It sits in Media & Creative, covering Transcription, Speech recognition and synthesis and REST APIs. It works with Azure AI Speech, Microsoft Azure and Python. The repository describes itself as: Skills, MCP servers, Custom Agents, Agents.md for SDKs to ground Coding Agents. The licence is MIT.

When your agent uses it

  • Simple speech recognition of audio files up to 60 seconds without the Speech SDK
  • : Long audio (60 seconds)
  • Real-time streaming
  • Batch transcription

Example prompts

  • “speech to text REST”
  • “short audio transcription”
  • “speech recognition REST API”
  • “/azure-speech-to-text-rest-py”

Requirements

  • Python 3
  • A credential in AZURE_SPEECH_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Azure subscription - Create one free
  2. Speech resource - Create in Azure Portal
  3. Get credentials - After deployment, go to resource > Keys and Endpoint

What it can do on your machine

Read from SKILL.md and the folder at commit d5741a1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • azure.microsoft.com
    • portal.azure.com
    • learn.microsoft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_SPEECH_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Azure Speech To Text REST Py loads about 3k tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 120 tokens; SKILL.md has 489 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/skills at commit d5741a1, republished under its MIT licence (© microsoft). 489 words, ~3,013 tokens.

Download SKILL.mdSave it as .claude/skills/azure-speech-to-text-rest-py/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
azure-speech-to-text-rest-py
description
Azure Speech to Text REST API for short audio (Python). Use for simple speech recognition of audio files up to 60 seconds without the Speech SDK. Triggers: "speech to text REST", "short audio transcription", "speech recognition REST API", "STT REST", "recognize speech REST". DO NOT USE FOR: Long audio (>60 seconds), real-time streaming, batch transcription, custom speech models, speech translation. Use Speech SDK or Batch Transcription API instead.
license
MIT
metadata.author
Microsoft
metadata.version
1.0.0

Azure Speech to Text REST API for Short Audio

Simple REST API for speech-to-text transcription of short audio files (up to 60 seconds). No SDK required - just HTTP requests.

Prerequisites

  1. Azure subscription - Create one free
  2. Speech resource - Create in Azure Portal
  3. Get credentials - After deployment, go to resource > Keys and Endpoint

Environment Variables

bash
# Required
AZURE_SPEECH_KEY=<your-speech-resource-key>
AZURE_SPEECH_REGION=<region>  # e.g., eastus, westus2, westeurope

# Alternative: Use endpoint directly
AZURE_SPEECH_ENDPOINT=https://<region>.stt.speech.microsoft.com

Installation

bash
pip install requests

Authentication & Lifecycle

🔑 Two rules apply to every code sample below:

  1. Two auth modes are supported. Use a subscription key (Ocp-Apim-Subscription-Key header) for quick access, or a Microsoft Entra token (including one acquired with DefaultAzureCredential) via the Authorization request header (see "Option 2" below). Never hardcode credentials in source.
  2. Use context managers for files and HTTP resources so file handles and network connections are released deterministically:
    • Sync: with open(...) as f: and (when reusing connections) with requests.Session() as session:
    • Async: async with aiohttp.ClientSession() as session:

Snippets may abbreviate this setup, but production code should always follow both rules.

Quick Start

python
import os
import requests

def transcribe_audio(audio_file_path: str, language: str = "en-US") -> dict:
    """Transcribe short audio file (max 60 seconds) using REST API."""
    region = os.environ["AZURE_SPEECH_REGION"]
    api_key = os.environ["AZURE_SPEECH_KEY"]
    
    url = f"https://{region}.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1"
    
    headers = {
        "Ocp-Apim-Subscription-Key": api_key,
        "Content-Type": "audio/wav; codecs=audio/pcm; samplerate=16000",
        "Accept": "application/json"
    }
    
    params = {
        "language": language,
        "format": "detailed"  # or "simple"
    }
    
    with open(audio_file_path, "rb") as audio_file:
        response = requests.post(url, headers=headers, params=params, data=audio_file)
    
    response.raise_for_status()
    return response.json()

# Usage
result = transcribe_audio("audio.wav", "en-US")
print(result["DisplayText"])

Audio Requirements

FormatCodecSample RateNotes
WAVPCM16 kHz, monoRecommended
OGGOPUS16 kHz, monoSmaller file size

Limitations:

  • Maximum 60 seconds of audio
  • For pronunciation assessment: maximum 30 seconds
  • No partial/interim results (final only)

Content-Type Headers

python
# WAV PCM 16kHz
wav_headers = {
    "Content-Type": "audio/wav; codecs=audio/pcm; samplerate=16000"
}

# OGG OPUS
ogg_headers = {
    "Content-Type": "audio/ogg; codecs=opus"
}

Response Formats

Simple Format (default)
python
params = {"language": "en-US", "format": "simple"}
json
{
  "RecognitionStatus": "Success",
  "DisplayText": "Remind me to buy 5 pencils.",
  "Offset": "1236645672289",
  "Duration": "1236645672289"
}
Detailed Format
python
params = {"language": "en-US", "format": "detailed"}
json
{
  "RecognitionStatus": "Success",
  "Offset": "1236645672289",
  "Duration": "1236645672289",
  "NBest": [
    {
      "Confidence": 0.9052885,
      "Display": "What's the weather like?",
      "ITN": "what's the weather like",
      "Lexical": "what's the weather like",
      "MaskedITN": "what's the weather like"
    }
  ]
}

For lower latency, stream audio in chunks:

python
import os
import requests

def transcribe_chunked(audio_file_path: str, language: str = "en-US") -> dict:
    """Stream audio in chunks for lower latency."""
    region = os.environ["AZURE_SPEECH_REGION"]
    api_key = os.environ["AZURE_SPEECH_KEY"]
    
    url = f"https://{region}.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1"
    
    headers = {
        "Ocp-Apim-Subscription-Key": api_key,
        "Content-Type": "audio/wav; codecs=audio/pcm; samplerate=16000",
        "Accept": "application/json",
        "Transfer-Encoding": "chunked",
        "Expect": "100-continue"
    }
    
    params = {"language": language, "format": "detailed"}
    
    def generate_chunks(file_path: str, chunk_size: int = 1024):
        with open(file_path, "rb") as f:
            while chunk := f.read(chunk_size):
                yield chunk
    
    response = requests.post(
        url, 
        headers=headers, 
        params=params, 
        data=generate_chunks(audio_file_path)
    )
    
    response.raise_for_status()
    return response.json()

Authentication Options

Option 1: Subscription Key (Simple)
python
headers = {
    "Ocp-Apim-Subscription-Key": os.environ["AZURE_SPEECH_KEY"]
}
Option 2: Bearer Token
python
import requests
import os

def get_access_token() -> str:
    """Get access token from the token endpoint."""
    region = os.environ["AZURE_SPEECH_REGION"]
    api_key = os.environ["AZURE_SPEECH_KEY"]
    
    token_url = f"https://{region}.api.cognitive.microsoft.com/sts/v1.0/issueToken"
    
    response = requests.post(
        token_url,
        headers={
            "Ocp-Apim-Subscription-Key": api_key,
            "Content-Type": "application/x-www-form-urlencoded",
            "Content-Length": "0"
        }
    )
    response.raise_for_status()
    return response.text

# Use token in requests (valid for 10 minutes)
token = get_access_token()
headers = {
    "Authorization": f"Bearer {token}",
    "Content-Type": "audio/wav; codecs=audio/pcm; samplerate=16000",
    "Accept": "application/json"
}

Query Parameters

ParameterRequiredValuesDescription
languageYesen-US, de-DE, etc.Language of speech
formatNosimple, detailedResult format (default: simple)
profanityNomasked, removed, rawProfanity handling (default: masked)

Recognition Status Values

StatusDescription
SuccessRecognition succeeded
NoMatchSpeech detected but no words matched
InitialSilenceTimeoutOnly silence detected
BabbleTimeoutOnly noise detected
ErrorInternal service error

Profanity Handling

python
# Mask profanity with asterisks (default)
params = {"language": "en-US", "profanity": "masked"}

# Remove profanity entirely
params = {"language": "en-US", "profanity": "removed"}

# Include profanity as-is
params = {"language": "en-US", "profanity": "raw"}

Error Handling

python
import requests

def transcribe_with_error_handling(audio_path: str, language: str = "en-US") -> dict | None:
    """Transcribe with proper error handling."""
    region = os.environ["AZURE_SPEECH_REGION"]
    api_key = os.environ["AZURE_SPEECH_KEY"]
    
    url = f"https://{region}.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1"
    
    try:
        with open(audio_path, "rb") as audio_file:
            response = requests.post(
                url,
                headers={
                    "Ocp-Apim-Subscription-Key": api_key,
                    "Content-Type": "audio/wav; codecs=audio/pcm; samplerate=16000",
                    "Accept": "application/json"
                },
                params={"language": language, "format": "detailed"},
                data=audio_file
            )
        
        if response.status_code == 200:
            result = response.json()
            if result.get("RecognitionStatus") == "Success":
                return result
            else:
                print(f"Recognition failed: {result.get('RecognitionStatus')}")
                return None
        elif response.status_code == 400:
            print(f"Bad request: Check language code or audio format")
        elif response.status_code == 401:
            print(f"Unauthorized: Check API key or token")
        elif response.status_code == 403:
            print(f"Forbidden: Missing authorization header")
        else:
            print(f"Error {response.status_code}: {response.text}")
        
        return None
        
    except requests.exceptions.RequestException as e:
        print(f"Request failed: {e}")
        return None

Async Version

python
import os
import aiohttp
import asyncio

async def transcribe_async(audio_file_path: str, language: str = "en-US") -> dict:
    """Async version using aiohttp."""
    region = os.environ["AZURE_SPEECH_REGION"]
    api_key = os.environ["AZURE_SPEECH_KEY"]
    
    url = f"https://{region}.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1"
    
    headers = {
        "Ocp-Apim-Subscription-Key": api_key,
        "Content-Type": "audio/wav; codecs=audio/pcm; samplerate=16000",
        "Accept": "application/json"
    }
    
    params = {"language": language, "format": "detailed"}
    
    async with aiohttp.ClientSession() as session:
        with open(audio_file_path, "rb") as f:
            audio_data = f.read()
        
        async with session.post(url, headers=headers, params=params, data=audio_data) as response:
            response.raise_for_status()
            return await response.json()

# Usage
result = asyncio.run(transcribe_async("audio.wav", "en-US"))
print(result["DisplayText"])
Show full SKILL.md (199 more words)Show less

Supported Languages

Common language codes (see full list):

CodeLanguage
en-USEnglish (US)
en-GBEnglish (UK)
de-DEGerman
fr-FRFrench
es-ESSpanish (Spain)
es-MXSpanish (Mexico)
zh-CNChinese (Mandarin)
ja-JPJapanese
ko-KRKorean
pt-BRPortuguese (Brazil)

Best Practices

  1. Pick sync OR async and stay consistent. Do not mix azure.xxx sync clients with azure.xxx.aio async clients in the same call path. Choose one mode per module.
  2. Use context managers for files and HTTP resources. Use with open(...) as f: and (when reusing connections) with requests.Session() as session: for sync code, or async with aiohttp.ClientSession() as session: for async code.
  3. Use WAV PCM 16kHz mono for best compatibility
  4. Enable chunked transfer for lower latency
  5. Cache access tokens for 9 minutes (valid for 10)
  6. Specify the correct language for accurate recognition
  7. Use detailed format when you need confidence scores
  8. Handle all RecognitionStatus values in production code

When NOT to Use This API

Use the Speech SDK or Batch Transcription API instead when you need:

  • Audio longer than 60 seconds
  • Real-time streaming transcription
  • Partial/interim results
  • Speech translation
  • Custom speech models
  • Batch transcription of many files

Reference Files

FileContents
references/pronunciation-assessment.mdPronunciation assessment parameters and scoring

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in .github/plugins/azure-sdk-python/skills/azure-speech-to-text-rest-py of microsoft/skills.

  • SKILL.md
  • references/pronunciation-assessment.md

Open the folder on GitHubat commit d5741a1

Used in 5 other repositories

We found 14 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in microsoft/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Azure Speech To Text REST Py next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure Speech To Text REST Py compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure Speech To Text REST Py this skillmicrosoft/skills3.1k5 repos~3kAutomated safety check: PassMIT
Watchmathiaschu/watch142—~4kAutomated safety check: WarnMIT
Speech Engineelevenlabs/skills482—~2.5kAutomated safety check: WarnMIT
Azure AImicrosoft/GitHub-Copilot-for-Azure2551 repos~852Automated safety check: PassMIT
Douyin DownloaderOpenMinis/MinisSkills446—~1.4kAutomated safety check: PassMIT
Watch Video Q&Abradautomates/claude-video18k—~4.3kAutomated safety check: NotesMIT

Similar skills

  • Watch

    mathiaschu/watch

    Watch a video from YouTube, Instagram, X/Twitter, Vimeo, TikTok or any of ~1800 yt-dlp sites (or a local path).

    142 GitHub stars~4k tokensUpdated 4 mo ago
    Media & CreativeAuto-check: warnings
  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • Azure AI

    microsoft/GitHub-Copilot-for-Azure

    Official

    A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.

    255 GitHub starsUsed in 1 repo~852 tokens
    Media & CreativeAuto-check passed
  • Douyin Downloader

    OpenMinis/MinisSkills

    Download Douyin (TikTok) videos from share links. An agent skill from OpenMinis/MinisSkills.

    446 GitHub stars~1.4k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Watch Video Q&A

    bradautomates/claude-video

    Lets the agent answer questions about a video from a URL or local file by downloading it, extracting frames and a transcript, or by sending it to Gemini's video model.

    18k GitHub stars~4.3k tokensUpdated 16 days ago
    Media & CreativeAuto-check: notes
  • Claude Real Video

    HUANGCHIHHUNGLeo/claude-real-video

    Watch a video for the user. An agent skill from HUANGCHIHHUNGLeo/claude-real-video.

    2.2k GitHub stars~639 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed

More from microsoft/skills

All 150 skills in this repo
  • Official

    Covers producer, consumer, and checkpoint-store setup for Azure Event Hubs streaming in Python, with Entra ID auth and partition targeting.

    3.1k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Auto-check passed
  • Frontend UI Dark TS

    microsoft/skills

    Official

    Build dark-themed React applications using Tailwind CSS with custom theming, glassmorphism effects, and Framer Motion animations.

    3.1k GitHub starsUsed in 5 repos~3.6k tokens
    Auto-check passed
  • Pydantic Models Py

    microsoft/skills

    Official

    Create Pydantic models following the multi-model pattern with Base, Create, Update, Response, and InDB variants.

    3.1k GitHub starsUsed in 5 repos~496 tokens
    Auto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Skill Creator

    microsoft/skills

    Official

    Guide for creating effective skills for AI coding agents working with Azure SDKs and Microsoft Foundry services.

    3.1k GitHub starsUsed in 5 repos~17k tokens
    Auto-check passed

Questions about Azure Speech To Text REST Py

What does Azure Speech To Text REST Py do?

Azure Speech to Text REST API for short audio (Python). An agent skill from microsoft/skills. Azure Speech To Text REST Py is an agent skill from microsoft/skills, published by the product's own GitHub organization. Azure Speech to Text REST API for short audio (Python).

When should I use Azure Speech To Text REST Py?

Azure Speech To Text REST Py fits situations like: simple speech recognition of audio files up to 60 seconds without the Speech SDK; : Long audio (60 seconds); real-time streaming; batch transcription.

How do I install Azure Speech To Text REST Py in Claude Code?

Run `npx skills add microsoft/skills --skill azure-speech-to-text-rest-py -a claude-code`. Or copy the skill folder (.github/plugins/azure-sdk-python/skills/azure-speech-to-text-rest-py in microsoft/skills) into .claude/skills/azure-speech-to-text-rest-py in your project. Claude Code loads it when a task matches its description.

How do I install Azure Speech To Text REST Py in Codex?

Run `npx skills add microsoft/skills --skill azure-speech-to-text-rest-py -a codex`. Or copy the skill folder (.github/plugins/azure-sdk-python/skills/azure-speech-to-text-rest-py in microsoft/skills) into .agents/skills/azure-speech-to-text-rest-py in your project. Codex loads it when a task matches its description.

Can I use Azure Speech To Text REST Py in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/skills --skill azure-speech-to-text-rest-py -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azure-speech-to-text-rest-py, .gemini/skills/azure-speech-to-text-rest-py, .github/skills/azure-speech-to-text-rest-py and .opencode/skills/azure-speech-to-text-rest-py in your project.

What does Azure Speech To Text REST Py need to run?

Going by SKILL.md and its folder, Azure Speech To Text REST Py needs the command-line tools its instructions call (pip) and credentials named AZURE_SPEECH_KEY. Our summary lists: Python 3; A credential in AZURE_SPEECH_KEY.

Does Azure Speech To Text REST Py access the network?

SKILL.md names 3 domains. As links in the text: azure.microsoft.com, portal.azure.com and learn.microsoft.com. This is read from the text; nothing was executed.

Is Azure Speech To Text REST Py safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Azure Speech To Text REST Py use?

Azure Speech To Text REST Py is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure Speech To Text REST Py use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.9k tokens, read only when the agent opens those files.

What are the alternatives to Azure Speech To Text REST Py?

Skills that share tags, products or a category with Azure Speech To Text REST Py: Watch (mathiaschu/watch, 142 stars), Speech Engine (elevenlabs/skills, 482 stars), Azure AI (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Douyin Downloader (OpenMinis/MinisSkills, 446 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure Speech To Text REST Py?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/skills, which has 3,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: microsoft/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.