Official agent skill

Azure AI Voicelive Py

by microsoft in microsoft/skills

Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).

OfficialMITAuto-check passedBackend & APIs

Install Azure AI Voicelive Py

skills CLI
$ npx skills add microsoft/skills --skill azure-ai-voicelive-py -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/skills azure-ai-voicelive-py --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/plugins/azure-sdk-python/skills/azure-ai-voicelive-py .claude/skills/azure-ai-voicelive-py && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azure-ai-voicelive-py
GitHub stars
3.1k
Token cost
~2.9k tokens
SKILL.md length
388 words
Files
4 (incl. references)
Skills in repo
150
Repo updated
First seen
Licence
MIT

At a glance

Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).

  • Works in 2 steps: This SDK is async-only; use the .aio… → Always use context managers for clients…
  • Creating Python applications that need real-time bidirectional audio communication with Azure AI
  • SKILL.md covers Installation, Environment Variables, Authentication & Lifecycle and Quick Start, plus 11 more sections
  • Calls pip; reaches cognitiveservices.azure.com and learn.microsoft.com; needs AZURE_TOKEN_CREDENTIALS and AZURE_COGNITIVE_SERVICES_KEY

What it does

Azure AI Voicelive Py is an agent skill from microsoft/skills, published by the product's own GitHub organization. Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/api-reference.md`, `references/examples.md` and `references/models.md`).

It sits in Backend & APIs, covering Realtime and WebSockets, Transcription and Speech recognition and synthesis. It works with Microsoft Azure, Azure AI Speech, Visual Studio Code and Python. The repository describes itself as: Skills, MCP servers, Custom Agents, Agents.md for SDKs to ground Coding Agents. The licence is MIT.

When your agent uses it

  • Creating Python applications that need real-time bidirectional audio communication with Azure AI
  • Including voice assistants
  • Voice-enabled chatbots
  • Real-time speech-to-speech translation

Example prompts

  • “/azure-ai-voicelive-py”

Requirements

  • Python 3
  • A credential in AZURE_COGNITIVE_SERVICES_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. This SDK is async-only; use the .aio namespace throughout. Do not try to pair it with sync clients from other Azure SDKs in the same call…
  2. Always use context managers for clients and async credentials. Wrap every connection in async with connect(...) as conn:. For async…

What it can do on your machine

Read from SKILL.md and the folder at commit d5741a1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • cognitiveservices.azure.com
    • learn.microsoft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_TOKEN_CREDENTIALS
    • AZURE_COGNITIVE_SERVICES_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Azure AI Voicelive Py loads about 2.9k tokens when it runs, and up to ~14k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 388 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~14k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/skills at commit d5741a1, republished under its MIT licence (© microsoft). 388 words, ~2,895 tokens.

Download SKILL.mdSave it as .claude/skills/azure-ai-voicelive-py/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
azure-ai-voicelive-py
description
Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Use this skill when creating Python applications that need real-time bidirectional audio communication with Azure AI, including voice assistants, voice-enabled chatbots, real-time speech-to-speech translation, voice-driven avatars, or any WebSocket-based audio streaming with AI models. Supports Server VAD (Voice Activity Detection), turn-based conversation, function calling, MCP tools, avatar integration, and transcription.
license
MIT
metadata.author
Microsoft
metadata.version
1.0.0
metadata.package
azure-ai-voicelive

Azure AI Voice Live SDK

Build real-time voice AI applications with bidirectional WebSocket communication.

Installation

bash
pip install azure-ai-voicelive aiohttp azure-identity

Environment Variables

bash
AZURE_COGNITIVE_SERVICES_ENDPOINT=https://<region>.api.cognitive.microsoft.com  # Required for all auth methods
AZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production
AZURE_COGNITIVE_SERVICES_KEY=<api-key>  # Only required for the legacy API-key auth path below

Authentication & Lifecycle

🔑 Two rules apply to every code sample below:

  1. Prefer DefaultAzureCredential. It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation.
    • Local dev: DefaultAzureCredential works as-is.
    • Production: set AZURE_TOKEN_CREDENTIALS=prod (or AZURE_TOKEN_CREDENTIALS=<specific_credential>) to constrain the credential chain to production-safe credentials.
  2. Wrap every client in a context manager so HTTP transports, sockets, and token caches are released deterministically:
    • Sync: with <Client>(...) as client:
    • Async: async with <Client>(...) as client: and async with DefaultAzureCredential() as credential: (from azure.identity.aio)

Snippets may abbreviate this setup, but production code should always follow both rules.

python
import os
from azure.ai.voicelive.aio import connect
from azure.identity.aio import DefaultAzureCredential, ManagedIdentityCredential

# Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential>
# Or use a specific credential directly in production:
# See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes
# credential = ManagedIdentityCredential()

async with DefaultAzureCredential(require_envvar=True) as credential:
    async with connect(
        endpoint=os.environ["AZURE_COGNITIVE_SERVICES_ENDPOINT"],
        credential=credential,
        model="gpt-4o-realtime-preview",
        credential_scopes=["https://cognitiveservices.azure.com/.default"]
    ) as conn:
        ...
Legacy: API Key (existing keyed deployments)

New code should use DefaultAzureCredential above. Use AzureKeyCredential only if you have an existing keyed deployment that hasn't been migrated to Entra ID yet — for example, regulated environments still completing their Entra rollout.

python
import os
from azure.core.credentials import AzureKeyCredential
from azure.ai.voicelive.aio import connect

async with connect(
    endpoint=os.environ["AZURE_COGNITIVE_SERVICES_ENDPOINT"],
    credential=AzureKeyCredential(os.environ["AZURE_COGNITIVE_SERVICES_KEY"]),
    model="gpt-4o-realtime-preview",
) as conn:
    ...

Quick Start

python
import asyncio
import os
from azure.ai.voicelive.aio import connect
from azure.identity.aio import DefaultAzureCredential

async def main():
    async with connect(
        endpoint=os.environ["AZURE_COGNITIVE_SERVICES_ENDPOINT"],
        credential=DefaultAzureCredential(),
        model="gpt-4o-realtime-preview",
        credential_scopes=["https://cognitiveservices.azure.com/.default"]
    ) as conn:
        # Update session with instructions
        await conn.session.update(session={
            "instructions": "You are a helpful assistant.",
            "modalities": ["text", "audio"],
            "voice": "alloy"
        })
        
        # Listen for events
        async for event in conn:
            print(f"Event: {event.type}")
            if event.type == "response.audio_transcript.done":
                print(f"Transcript: {event.transcript}")
            elif event.type == "response.done":
                break

asyncio.run(main())

Core Architecture

Connection Resources

The VoiceLiveConnection exposes these resources:

ResourcePurposeKey Methods
conn.sessionSession configurationupdate(session=...)
conn.responseModel responsescreate(), cancel()
conn.input_audio_bufferAudio inputappend(), commit(), clear()
conn.output_audio_bufferAudio outputclear()
conn.conversationConversation stateitem.create(), item.delete(), item.truncate()
conn.transcription_sessionTranscription configupdate(session=...)

Session Configuration

python
from azure.ai.voicelive.models import RequestSession, FunctionTool

await conn.session.update(session=RequestSession(
    instructions="You are a helpful voice assistant.",
    modalities=["text", "audio"],
    voice="alloy",  # or "echo", "shimmer", "sage", etc.
    input_audio_format="pcm16",
    output_audio_format="pcm16",
    turn_detection={
        "type": "server_vad",
        "threshold": 0.5,
        "prefix_padding_ms": 300,
        "silence_duration_ms": 500
    },
    tools=[
        FunctionTool(
            type="function",
            name="get_weather",
            description="Get current weather",
            parameters={
                "type": "object",
                "properties": {
                    "location": {"type": "string"}
                },
                "required": ["location"]
            }
        )
    ]
))

Audio Streaming

Send Audio (Base64 PCM16)
python
import base64

# Read audio chunk (16-bit PCM, 24kHz mono)
audio_chunk = await read_audio_from_microphone()
b64_audio = base64.b64encode(audio_chunk).decode()

await conn.input_audio_buffer.append(audio=b64_audio)
Receive Audio
python
async for event in conn:
    if event.type == "response.audio.delta":
        audio_bytes = base64.b64decode(event.delta)
        await play_audio(audio_bytes)
    elif event.type == "response.audio.done":
        print("Audio complete")

Event Handling

python
async for event in conn:
    match event.type:
        # Session events
        case "session.created":
            print(f"Session: {event.session}")
        case "session.updated":
            print("Session updated")
        
        # Audio input events
        case "input_audio_buffer.speech_started":
            print(f"Speech started at {event.audio_start_ms}ms")
        case "input_audio_buffer.speech_stopped":
            print(f"Speech stopped at {event.audio_end_ms}ms")
        
        # Transcription events
        case "conversation.item.input_audio_transcription.completed":
            print(f"User said: {event.transcript}")
        case "conversation.item.input_audio_transcription.delta":
            print(f"Partial: {event.delta}")
        
        # Response events
        case "response.created":
            print(f"Response started: {event.response.id}")
        case "response.audio_transcript.delta":
            print(event.delta, end="", flush=True)
        case "response.audio.delta":
            audio = base64.b64decode(event.delta)
        case "response.done":
            print(f"Response complete: {event.response.status}")
        
        # Function calls
        case "response.function_call_arguments.done":
            result = handle_function(event.name, event.arguments)
            await conn.conversation.item.create(item={
                "type": "function_call_output",
                "call_id": event.call_id,
                "output": json.dumps(result)
            })
            await conn.response.create()
        
        # Errors
        case "error":
            print(f"Error: {event.error.message}")

Common Patterns

Manual Turn Mode (No VAD)
python
await conn.session.update(session={"turn_detection": None})

# Manually control turns
await conn.input_audio_buffer.append(audio=b64_audio)
await conn.input_audio_buffer.commit()  # End of user turn
await conn.response.create()  # Trigger response
Interrupt Handling
python
async for event in conn:
    if event.type == "input_audio_buffer.speech_started":
        # User interrupted - cancel current response
        await conn.response.cancel()
        await conn.output_audio_buffer.clear()
Show full SKILL.md (155 more words)Show less
Conversation History
python
# Add system message
await conn.conversation.item.create(item={
    "type": "message",
    "role": "system",
    "content": [{"type": "input_text", "text": "Be concise."}]
})

# Add user message
await conn.conversation.item.create(item={
    "type": "message",
    "role": "user", 
    "content": [{"type": "input_text", "text": "Hello!"}]
})

await conn.response.create()

Voice Options

VoiceDescription
alloyNeutral, balanced
echoWarm, conversational
shimmerClear, professional
sageCalm, authoritative
coralFriendly, upbeat
ashDeep, measured
balladExpressive
verseStorytelling

Azure voices: Use AzureStandardVoice, AzureCustomVoice, or AzurePersonalVoice models.

Audio Formats

FormatSample RateUse Case
pcm1624kHzDefault, high quality
pcm16-8000hz8kHzTelephony
pcm16-16000hz16kHzVoice assistants
g711_ulaw8kHzTelephony (US)
g711_alaw8kHzTelephony (EU)

Turn Detection Options

python
# Server VAD (default)
{"type": "server_vad", "threshold": 0.5, "silence_duration_ms": 500}

# Azure Semantic VAD (smarter detection)
{"type": "azure_semantic_vad"}
{"type": "azure_semantic_vad_en"}  # English optimized
{"type": "azure_semantic_vad_multilingual"}

Error Handling

python
from azure.ai.voicelive.aio import ConnectionError, ConnectionClosed

try:
    async with connect(...) as conn:
        async for event in conn:
            if event.type == "error":
                print(f"API Error: {event.error.code} - {event.error.message}")
except ConnectionClosed as e:
    print(f"Connection closed: {e.code} - {e.reason}")
except ConnectionError as e:
    print(f"Connection error: {e}")

Best Practices

  1. This SDK is async-only; use the .aio namespace throughout. Do not try to pair it with sync clients from other Azure SDKs in the same call path — keep the whole request path async.
  2. Always use context managers for clients and async credentials. Wrap every connection in async with connect(...) as conn:. For async DefaultAzureCredential from azure.identity.aio, also use async with credential: so tokens and transports are cleaned up.

References

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .github/plugins/azure-sdk-python/skills/azure-ai-voicelive-py of microsoft/skills.

  • SKILL.md
  • references/api-reference.md
  • references/examples.md
  • references/models.md

Open the folder on GitHubat commit d5741a1

Compare with similar skills

Azure AI Voicelive Py next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure AI Voicelive Py compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure AI Voicelive Py this skillmicrosoft/skills3.1k—~2.9kAutomated safety check: PassMIT
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
Deepgram Python Speech-to-Textdeepgram/deepgram-python-sdk469—~2.9kAutomated safety check: PassMIT
Deepgram Python Voice Agentdeepgram/deepgram-python-sdk469—~3.6kAutomated safety check: PassMIT
Speech Engineelevenlabs/skills482—~2.5kAutomated safety check: WarnMIT
Deepgram JS Audio Intelligencedeepgram/deepgram-js-sdk276—~1.5kAutomated safety check: PassMIT

Similar skills

  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed
  • Deepgram Python Speech-to-Text

    deepgram/deepgram-python-sdk

    Covers basic transcription with the Deepgram Python SDK's listen.v1 endpoint, for one-shot REST transcription of a file or URL and live WebSocket streaming with interim results.

    469 GitHub stars~2.9k tokensUpdated 2 days ago
    Backend & APIsAuto-check passed
  • Deepgram Python Voice Agent

    deepgram/deepgram-python-sdk

    Builds a full-duplex Python voice agent on Deepgram's agent.converse WebSocket, combining speech-to-text, an LLM and text-to-speech with interruption and function calling.

    469 GitHub stars~3.6k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • Deepgram JS Audio Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…

    276 GitHub stars~1.5k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Gemini Live API Dev

    JetBrains/skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API.

    366 GitHub stars~2.6k tokensUpdated 3 mo ago
    Backend & APIsAuto-check passed

More from microsoft/skills

All 150 skills in this repo
  • Official

    Covers producer, consumer, and checkpoint-store setup for Azure Event Hubs streaming in Python, with Entra ID auth and partition targeting.

    3.1k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Auto-check passed
  • Frontend UI Dark TS

    microsoft/skills

    Official

    Build dark-themed React applications using Tailwind CSS with custom theming, glassmorphism effects, and Framer Motion animations.

    3.1k GitHub starsUsed in 5 repos~3.6k tokens
    Auto-check passed
  • Pydantic Models Py

    microsoft/skills

    Official

    Create Pydantic models following the multi-model pattern with Base, Create, Update, Response, and InDB variants.

    3.1k GitHub starsUsed in 5 repos~496 tokens
    Auto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub stars~2.8k tokensUpdated yesterday
    Auto-check passed
  • Skill Creator

    microsoft/skills

    Official

    Guide for creating effective skills for AI coding agents working with Azure SDKs and Microsoft Foundry services.

    3.1k GitHub starsUsed in 5 repos~17k tokens
    Auto-check passed

Questions about Azure AI Voicelive Py

What does Azure AI Voicelive Py do?

Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive). Azure AI Voicelive Py is an agent skill from microsoft/skills, published by the product's own GitHub organization. Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).

When should I use Azure AI Voicelive Py?

Azure AI Voicelive Py fits situations like: creating Python applications that need real-time bidirectional audio communication with Azure AI; including voice assistants; voice-enabled chatbots; real-time speech-to-speech translation.

How do I install Azure AI Voicelive Py in Claude Code?

Run `npx skills add microsoft/skills --skill azure-ai-voicelive-py -a claude-code`. Or copy the skill folder (.github/plugins/azure-sdk-python/skills/azure-ai-voicelive-py in microsoft/skills) into .claude/skills/azure-ai-voicelive-py in your project. Claude Code loads it when a task matches its description.

How do I install Azure AI Voicelive Py in Codex?

Run `npx skills add microsoft/skills --skill azure-ai-voicelive-py -a codex`. Or copy the skill folder (.github/plugins/azure-sdk-python/skills/azure-ai-voicelive-py in microsoft/skills) into .agents/skills/azure-ai-voicelive-py in your project. Codex loads it when a task matches its description.

Can I use Azure AI Voicelive Py in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/skills --skill azure-ai-voicelive-py -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azure-ai-voicelive-py, .gemini/skills/azure-ai-voicelive-py, .github/skills/azure-ai-voicelive-py and .opencode/skills/azure-ai-voicelive-py in your project.

What does Azure AI Voicelive Py need to run?

Going by SKILL.md and its folder, Azure AI Voicelive Py needs the command-line tools its instructions call (pip) and credentials named AZURE_TOKEN_CREDENTIALS and AZURE_COGNITIVE_SERVICES_KEY. Our summary lists: Python 3; A credential in AZURE_COGNITIVE_SERVICES_KEY.

Does Azure AI Voicelive Py access the network?

SKILL.md names 2 domains. In commands or code: cognitiveservices.azure.com and learn.microsoft.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Azure AI Voicelive Py safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Azure AI Voicelive Py use?

Azure AI Voicelive Py is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure AI Voicelive Py use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 11k tokens, read only when the agent opens those files.

What are the alternatives to Azure AI Voicelive Py?

Skills that share tags, products or a category with Azure AI Voicelive Py: Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars), Deepgram Python Speech-to-Text (deepgram/deepgram-python-sdk, 469 stars), Deepgram Python Voice Agent (deepgram/deepgram-python-sdk, 469 stars) and Speech Engine (elevenlabs/skills, 482 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure AI Voicelive Py?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/skills, which has 3,097 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: microsoft/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.