Official agent skill

Azure Realtime Podcast Generation

by microsoft in microsoft/skills

Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

OfficialMITAuto-check passedMedia & Creative

Install Azure Realtime Podcast Generation

skills CLI
$ npx skills add microsoft/skills --skill podcast-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/skills podcast-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/podcast-generation .claude/skills/podcast-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
podcast-generation
GitHub stars
3.1k
Used in
1 other repo
Token cost
~947 tokens
SKILL.md length
150 words
Files
4 (incl. scripts, references)
Skills in repo
150
Repo updated
First seen
Licence
MIT

At a glance

Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

  • Works in 5 steps: Configure environment variables for… → Connect via WebSocket to Azure OpenAI… → Send text prompt, collect PCM audio… → …
  • Adding text-to-speech narration to a web app with the Azure OpenAI Realtime API
  • SKILL.md covers Quick Start, Environment Configuration, Core Workflow and Voice Options, plus 3 more sections
  • Runs Python scripts from its folder; needs AZURE_OPENAI_AUDIO_API_KEY

What it does

The skill lays out a full-stack pattern for turning text into spoken audio. A Python backend connects to the Azure OpenAI Realtime endpoint over WebSocket, sends a text prompt and collects PCM audio chunks along with a transcript. It converts the PCM (24kHz, 16-bit, mono) to WAV, base64-encodes it and returns it to the frontend, where React code turns it into a playable blob.

Configuration uses three environment variables for the audio API key, endpoint and deployment name, and the endpoint should be the bare resource URL without /openai/v1/. The skill lists six voices (alloy, echo, fable, onyx, nova, shimmer) and the Realtime events to handle for audio chunks, transcript text, completion and errors. Reference files cover the architecture and code examples, and scripts/pcm_to_wav.py does the audio conversion.

When your agent uses it

  • Adding text-to-speech narration to a web app with the Azure OpenAI Realtime API
  • Generating a podcast-style audio summary from written content
  • Converting raw PCM audio from the Realtime API into WAV

Example prompts

  • “Add an endpoint that turns an article into a podcast-style narration using GPT Realtime Mini.”
  • “Write the React component that plays back the base64 WAV from our backend.”
  • “Convert the PCM chunks we collect from the Realtime API into a WAV file.”

Requirements

  • An Azure OpenAI resource with a GPT Realtime Mini deployment
  • Python with FastAPI

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Configure environment variables for Realtime API
  2. Connect via WebSocket to Azure OpenAI Realtime endpoint
  3. Send text prompt, collect PCM audio chunks + transcript
  4. Convert PCM to WAV format
  5. Return base64-encoded audio to frontend for playback

What it can do on your machine

Read from SKILL.md and the folder at commit 354361d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_OPENAI_AUDIO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Azure Realtime Podcast Generation loads about 947 tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 101 tokens; SKILL.md has 150 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~947
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from microsoft/skills at commit 354361d, republished under its MIT licence (© microsoft). 150 words, ~947 tokens.

Download SKILL.mdSave it as .claude/skills/podcast-generation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
podcast-generation
description
Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket. Use when building text-to-speech features, audio narrative generation, podcast creation from content, or integrating with Azure OpenAI Realtime API for real audio output. Covers full-stack implementation from React frontend to Python FastAPI backend with WebSocket streaming.

Podcast Generation with GPT Realtime Mini

Generate real audio narratives from text content using Azure OpenAI's Realtime API.

Quick Start

  1. Configure environment variables for Realtime API
  2. Connect via WebSocket to Azure OpenAI Realtime endpoint
  3. Send text prompt, collect PCM audio chunks + transcript
  4. Convert PCM to WAV format
  5. Return base64-encoded audio to frontend for playback

Environment Configuration

env
AZURE_OPENAI_AUDIO_API_KEY=your_realtime_api_key
AZURE_OPENAI_AUDIO_ENDPOINT=https://your-resource.cognitiveservices.azure.com
AZURE_OPENAI_AUDIO_DEPLOYMENT=gpt-realtime-mini

Note: Endpoint should NOT include /openai/v1/ - just the base URL.

Core Workflow

Backend Audio Generation
python
from openai import AsyncOpenAI
import base64

# Convert HTTPS endpoint to WebSocket URL
ws_url = endpoint.replace("https://", "wss://") + "/openai/v1"

client = AsyncOpenAI(
    websocket_base_url=ws_url,
    api_key=api_key
)

audio_chunks = []
transcript_parts = []

async with client.realtime.connect(model="gpt-realtime-mini") as conn:
    # Configure for audio-only output
    await conn.session.update(session={
        "output_modalities": ["audio"],
        "instructions": "You are a narrator. Speak naturally."
    })
    
    # Send text to narrate
    await conn.conversation.item.create(item={
        "type": "message",
        "role": "user",
        "content": [{"type": "input_text", "text": prompt}]
    })
    
    await conn.response.create()
    
    # Collect streaming events
    async for event in conn:
        if event.type == "response.output_audio.delta":
            audio_chunks.append(base64.b64decode(event.delta))
        elif event.type == "response.output_audio_transcript.delta":
            transcript_parts.append(event.delta)
        elif event.type == "response.done":
            break

# Convert PCM to WAV (see scripts/pcm_to_wav.py)
pcm_audio = b''.join(audio_chunks)
wav_audio = pcm_to_wav(pcm_audio, sample_rate=24000)
Frontend Audio Playback
javascript
// Convert base64 WAV to playable blob
const base64ToBlob = (base64, mimeType) => {
  const bytes = atob(base64);
  const arr = new Uint8Array(bytes.length);
  for (let i = 0; i < bytes.length; i++) arr[i] = bytes.charCodeAt(i);
  return new Blob([arr], { type: mimeType });
};

const audioBlob = base64ToBlob(response.audio_data, 'audio/wav');
const audioUrl = URL.createObjectURL(audioBlob);
new Audio(audioUrl).play();

Voice Options

VoiceCharacter
alloyNeutral
echoWarm
fableExpressive
onyxDeep
novaFriendly
shimmerClear

Realtime API Events

  • response.output_audio.delta - Base64 audio chunk
  • response.output_audio_transcript.delta - Transcript text
  • response.done - Generation complete
  • error - Handle with event.error.message

Audio Format

  • Input: Text prompt
  • Output: PCM audio (24kHz, 16-bit, mono)
  • Storage: Base64-encoded WAV

References

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts, references) in .github/skills/podcast-generation of microsoft/skills.

  • SKILL.md
  • references/architecture.md
  • references/code-examples.md
  • scripts/pcm_to_wav.py

Open the folder on GitHubat commit 354361d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in microsoft/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Azure Realtime Podcast Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure Realtime Podcast Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure Realtime Podcast Generation this skillmicrosoft/skills3.1k1 repos~947Automated safety check: PassMIT
Starlettesimonw/research783—~8.6kAutomated safety check: NotesNone
Fastapi Prodavila7/claude-code-templates32k8 repos~1.6kAutomated safety check: PassMIT
FastAPI ExpertJeffallan/claude-skills12k—~1.8kAutomated safety check: PassMIT
Fullstack DevHHU3637kr/skills1453 repos~8.6kAutomated safety check: NotesMIT
Deploy Fullstack Vercelvellum-ai/vellum-assistant1.4k—~2.8kAutomated safety check: PassMIT

Similar skills

  • Starlette

    simonw/research

    Build async web applications and APIs with Starlette 1.0, the lightweight ASGI framework for Python.

    783 GitHub stars~8.6k tokensUpdated 3 days ago
    Backend & APIsAuto-check: notes
  • Fastapi Pro

    davila7/claude-code-templates

    Build high-performance async APIs with FastAPI, SQLAlchemy 2.0, and Pydantic V2.

    32k GitHub starsUsed in 8 repos~1.6k tokens
    Backend & APIsAuto-check passed
  • FastAPI Expert

    Jeffallan/claude-skills

    Builds async Python APIs with FastAPI and Pydantic V2, covering endpoints, JWT authentication, async SQLAlchemy, WebSockets and pytest checks against the OpenAPI docs.

    12k GitHub stars~1.8k tokensUpdated 5 days ago
    Backend & APIsAuto-check passed
  • Fullstack Dev

    HHU3637kr/skills

    Full-stack backend architecture and frontend-backend integration guide.

    145 GitHub starsUsed in 3 repos~8.6k tokens
    Backend & APIsAuto-check: notes
  • Deploy Fullstack Vercel

    vellum-ai/vellum-assistant

    Build and deploy a full-stack app (React frontend + Python/FastAPI backend) or a Vellum app to Vercel as a serverless demo with seeded data

    1.4k GitHub stars~2.8k tokensUpdated today
    Backend & APIsAuto-check passed
  • Airflow Plugins

    astronomer/agents

    Builds Airflow 3.1+ plugins that embed FastAPI apps, custom UI pages, React components, middleware, macros, and operator links directly into the Airflow UI.

    451 GitHub stars~6k tokensUpdated yesterday
    Data & AnalyticsAuto-check: notes

More from microsoft/skills

All 150 skills in this repo
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub starsUsed in 6 repos~2.8k tokens
    Auto-check passed
  • Official

    Python guidance for the Azure AI Search SDK covering vector, hybrid and semantic search, index management and indexers, with Entra ID authentication preferred over keys.

    3.1k GitHub starsUsed in 6 repos~4.4k tokens
    Auto-check passed
  • Official

    Covers producer, consumer, and checkpoint-store setup for Azure Event Hubs streaming in Python, with Entra ID auth and partition targeting.

    3.1k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Pydantic Models Py

    microsoft/skills

    Official

    Create Pydantic models following the multi-model pattern with Base, Create, Update, Response, and InDB variants.

    3.1k GitHub starsUsed in 6 repos~496 tokens
    Auto-check passed
  • DebugView CLI

    microsoft/skills

    Official

    Captures and filters Windows user-mode and kernel debug output from the command line with the Sysinternals DebugView CLI, including bounded runs suited to agents.

    3.1k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Frontend UI Dark TS

    microsoft/skills

    Official

    Build dark-themed React applications using Tailwind CSS with custom theming, glassmorphism effects, and Framer Motion animations.

    3.1k GitHub starsUsed in 5 repos~3.6k tokens
    Auto-check passed

Questions about Azure Realtime Podcast Generation

What does Azure Realtime Podcast Generation do?

Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player. The skill lays out a full-stack pattern for turning text into spoken audio. A Python backend connects to the Azure OpenAI Realtime endpoint over WebSocket, sends a text prompt and collects PCM audio chunks along with a transcript.

When should I use Azure Realtime Podcast Generation?

Azure Realtime Podcast Generation fits situations like: adding text-to-speech narration to a web app with the Azure OpenAI Realtime API; generating a podcast-style audio summary from written content; converting raw PCM audio from the Realtime API into WAV.

How do I install Azure Realtime Podcast Generation in Claude Code?

Run `npx skills add microsoft/skills --skill podcast-generation -a claude-code`. Or copy the skill folder (.github/skills/podcast-generation in microsoft/skills) into .claude/skills/podcast-generation in your project. Claude Code loads it when a task matches its description.

How do I install Azure Realtime Podcast Generation in Codex?

Run `npx skills add microsoft/skills --skill podcast-generation -a codex`. Or copy the skill folder (.github/skills/podcast-generation in microsoft/skills) into .agents/skills/podcast-generation in your project. Codex loads it when a task matches its description.

Can I use Azure Realtime Podcast Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/skills --skill podcast-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/podcast-generation, .gemini/skills/podcast-generation, .github/skills/podcast-generation and .opencode/skills/podcast-generation in your project.

What does Azure Realtime Podcast Generation need to run?

Going by SKILL.md and its folder, Azure Realtime Podcast Generation needs Python for the scripts in its folder and credentials named AZURE_OPENAI_AUDIO_API_KEY. Our summary lists: An Azure OpenAI resource with a GPT Realtime Mini deployment; Python with FastAPI.

Does Azure Realtime Podcast Generation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Azure Realtime Podcast Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Azure Realtime Podcast Generation use?

Azure Realtime Podcast Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure Realtime Podcast Generation use?

About 947 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to Azure Realtime Podcast Generation?

Skills that share tags, products or a category with Azure Realtime Podcast Generation: Starlette (simonw/research, 783 stars), Fastapi Pro (davila7/claude-code-templates, 32k stars), FastAPI Expert (Jeffallan/claude-skills, 12k stars) and Fullstack Dev (HHU3637kr/skills, 145 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure Realtime Podcast Generation?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/skills, which has 3,091 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 6, 2026.

Source: microsoft/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.