Agent skill

Podcast Generation

by aAAaqwq in aAAaqwq/AGI-Super-Team

Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket.

MITAuto-check passedMedia & Creative

Install Podcast Generation

skills CLI
$ npx skills add aAAaqwq/AGI-Super-Team --skill podcast-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aAAaqwq/AGI-Super-Team podcast-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/podcast-generation .claude/skills/podcast-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
podcast-generation
GitHub stars
105
Used in
5 other repos
Token cost
~913 tokens
SKILL.md length
167 words
Files
1
Skills in repo
167
Repo updated
First seen
Licence
MIT

At a glance

Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket.

  • Works in 5 steps: Configure environment variables for… → Connect via WebSocket to Azure OpenAI… → Send text prompt, collect PCM audio… → …
  • Building text-to-speech features
  • SKILL.md covers Quick Start, Environment Configuration, Core Workflow and Voice Options, plus 4 more sections
  • Needs AZURE_OPENAI_AUDIO_API_KEY

What it does

Podcast Generation is an agent skill from aAAaqwq/AGI-Super-Team. Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket. Use when building text-to-speech features, audio narrative generation, podcast creatio...

Its SKILL.md is about 910 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Realtime and WebSockets and Text to speech and voice. It works with Azure OpenAI. The repository describes itself as: An installable, cross-framework AI organization: C-suite agents, expert subagents, curated skills, independent review, and one-command setup across 18 AI client/runtime adapters. The licence is MIT.

When your agent uses it

  • Building text-to-speech features
  • Audio narrative generation
  • Podcast creatio..

Example prompts

  • “/podcast-generation”

Requirements

  • Python 3
  • A credential in AZURE_OPENAI_AUDIO_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Configure environment variables for Realtime API
  2. Connect via WebSocket to Azure OpenAI Realtime endpoint
  3. Send text prompt, collect PCM audio chunks + transcript
  4. Convert PCM to WAV format
  5. Return base64-encoded audio to frontend for playback

What it can do on your machine

Read from SKILL.md and the folder at commit 7cefd81. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are env, python and javascript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_OPENAI_AUDIO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Podcast Generation loads about 913 tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 167 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~55
When it runs · the whole SKILL.md, loaded when a task matches
~913

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aAAaqwq/AGI-Super-Team at commit 7cefd81, republished under its MIT licence (© aAAaqwq). 167 words, ~913 tokens.

Download SKILL.mdSave it as .claude/skills/podcast-generation/SKILL.md (or your agent's skills folder).
name
podcast-generation
description
Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket. Use when building text-to-speech features, audio narrative generation, podcast creatio...
risk
unknown
source
community

Podcast Generation with GPT Realtime Mini

Generate real audio narratives from text content using Azure OpenAI's Realtime API.

Quick Start

  1. Configure environment variables for Realtime API
  2. Connect via WebSocket to Azure OpenAI Realtime endpoint
  3. Send text prompt, collect PCM audio chunks + transcript
  4. Convert PCM to WAV format
  5. Return base64-encoded audio to frontend for playback

Environment Configuration

env
AZURE_OPENAI_AUDIO_API_KEY=your_realtime_api_key
AZURE_OPENAI_AUDIO_ENDPOINT=https://your-resource.cognitiveservices.azure.com
AZURE_OPENAI_AUDIO_DEPLOYMENT=gpt-realtime-mini

Note: Endpoint should NOT include /openai/v1/ - just the base URL.

Core Workflow

Backend Audio Generation
python
from openai import AsyncOpenAI
import base64

# Convert HTTPS endpoint to WebSocket URL
ws_url = endpoint.replace("https://", "wss://") + "/openai/v1"

client = AsyncOpenAI(
    websocket_base_url=ws_url,
    api_key=api_key
)

audio_chunks = []
transcript_parts = []

async with client.realtime.connect(model="gpt-realtime-mini") as conn:
    # Configure for audio-only output
    await conn.session.update(session={
        "output_modalities": ["audio"],
        "instructions": "You are a narrator. Speak naturally."
    })
    
    # Send text to narrate
    await conn.conversation.item.create(item={
        "type": "message",
        "role": "user",
        "content": [{"type": "input_text", "text": prompt}]
    })
    
    await conn.response.create()
    
    # Collect streaming events
    async for event in conn:
        if event.type == "response.output_audio.delta":
            audio_chunks.append(base64.b64decode(event.delta))
        elif event.type == "response.output_audio_transcript.delta":
            transcript_parts.append(event.delta)
        elif event.type == "response.done":
            break

# Convert PCM to WAV (see scripts/pcm_to_wav.py)
pcm_audio = b''.join(audio_chunks)
wav_audio = pcm_to_wav(pcm_audio, sample_rate=24000)
Frontend Audio Playback
javascript
// Convert base64 WAV to playable blob
const base64ToBlob = (base64, mimeType) => {
  const bytes = atob(base64);
  const arr = new Uint8Array(bytes.length);
  for (let i = 0; i < bytes.length; i++) arr[i] = bytes.charCodeAt(i);
  return new Blob([arr], { type: mimeType });
};

const audioBlob = base64ToBlob(response.audio_data, 'audio/wav');
const audioUrl = URL.createObjectURL(audioBlob);
new Audio(audioUrl).play();

Voice Options

VoiceCharacter
alloyNeutral
echoWarm
fableExpressive
onyxDeep
novaFriendly
shimmerClear

Realtime API Events

  • response.output_audio.delta - Base64 audio chunk
  • response.output_audio_transcript.delta - Transcript text
  • response.done - Generation complete
  • error - Handle with event.error.message

Audio Format

  • Input: Text prompt
  • Output: PCM audio (24kHz, 16-bit, mono)
  • Storage: Base64-encoded WAV

References

  • Full architecture: See references/architecture.md for complete stack design
  • Code examples: See references/code-examples.md for production patterns
  • PCM conversion: Use scripts/pcm_to_wav.py for audio format conversion

When to Use

This skill is applicable to execute the workflow or actions described in the overview.

© aAAaqwq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/podcast-generation of aAAaqwq/AGI-Super-Team.

Open the folder on GitHubat commit 7cefd81

Used in 5 other repositories

We found 15 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 5 other GitHub owners. This page covers the copy in aAAaqwq/AGI-Super-Team, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Podcast Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Podcast Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Podcast Generation this skillaAAaqwq/AGI-Super-Team1055 repos~913Automated safety check: PassMIT
Azure Realtime Podcast Generationmicrosoft/skills3.1k1 repos~947Automated safety check: PassMIT
Deepgram Python Text-to-Speechdeepgram/deepgram-python-sdk469—~1.8kAutomated safety check: PassMIT
Speech Engineelevenlabs/skills482—~2.5kAutomated safety check: WarnMIT
Grok Realtime Voice Integrationcursor/plugins11k—~1.7kAutomated safety check: PassNone
Aliyun Qwen Tts Realtimecinience/alicloud-skills397—~788Automated safety check: PassMIT

Similar skills

  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Media & CreativeAuto-check passed
  • Deepgram Python Text-to-Speech

    deepgram/deepgram-python-sdk

    Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper.

    469 GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • Official

    Wires Grok speech-to-speech into an app's own microphone and audio playback over a realtime WebSocket, replacing an STT-LLM-TTS cascade or OpenAI Realtime.

    11k GitHub stars~1.7k tokensUpdated today
    Media & CreativeAuto-check passed
  • Aliyun Qwen Tts Realtime

    cinience/alicloud-skills

    A skill your agent uses when real-time speech synthesis is needed with Alibaba Cloud Model Studio Qwen TTS Realtime models.

    397 GitHub stars~788 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Aivisspeech Engine Sentry Triage

    Aivis-Project/AivisSpeech-Engine

    AivisSpeech Engine の Sentry issue を調査し、修正すべきエンジン側の不具合と、入力値・ローカル環境・外部サービス由来のノイズを切り分けるためのスキルです。Sentry 側で既知ノイズを永続アーカイブする作業や、voicevoxengine/utility/sentryutility.py と関連テストを更新して既知ノイズを送信前に破棄する作業で使用します。

    182 GitHub stars~545 tokensUpdated today
    Media & CreativeAuto-check passed

More from aAAaqwq/AGI-Super-Team

All 167 skills in this repo
  • Content Creator

    aAAaqwq/AGI-Super-Team

    Create SEO-optimized marketing content with consistent brand voice.

    105 GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • Financial Calculator

    aAAaqwq/AGI-Super-Team

    Advanced financial calculator with future value tables, present value, discount calculations, markup pricing, and compound interest.

    105 GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Bankr Signals

    aAAaqwq/AGI-Super-Team

    Transaction-verified trading signals on Base blockchain. An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub starsUsed in 2 repos~3.3k tokens
    Auto-check passed
  • Erc 8004

    aAAaqwq/AGI-Super-Team

    Register AI agents on Ethereum mainnet using ERC-8004 (Trustless Agents).

    105 GitHub starsUsed in 2 repos~1.2k tokens
    Auto-check passed
  • Frontend Design Ultimate

    aAAaqwq/AGI-Super-Team

    Create distinctive, production-grade static sites with React, Tailwind CSS, and shadcn/ui — no mockups needed.

    105 GitHub starsUsed in 2 repos~2.7k tokens
    Auto-check passed
  • Zsxq Smart Publish

    aAAaqwq/AGI-Super-Team

    Publish and manage content on 知识星球 (zsxq.com). An agent skill from aAAaqwq/AGI-Super-Team.

    105 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Podcast Generation

What does Podcast Generation do?

Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket. Podcast Generation is an agent skill from aAAaqwq/AGI-Super-Team. Generate AI-powered podcast-style audio narratives using Azure OpenAI's GPT Realtime Mini model via WebSocket.

When should I use Podcast Generation?

Podcast Generation fits situations like: building text-to-speech features; audio narrative generation; podcast creatio..

How do I install Podcast Generation in Claude Code?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill podcast-generation -a claude-code`. Or copy the skill folder (skills/podcast-generation in aAAaqwq/AGI-Super-Team) into .claude/skills/podcast-generation in your project. Claude Code loads it when a task matches its description.

How do I install Podcast Generation in Codex?

Run `npx skills add aAAaqwq/AGI-Super-Team --skill podcast-generation -a codex`. Or copy the skill folder (skills/podcast-generation in aAAaqwq/AGI-Super-Team) into .agents/skills/podcast-generation in your project. Codex loads it when a task matches its description.

Can I use Podcast Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aAAaqwq/AGI-Super-Team --skill podcast-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/podcast-generation, .gemini/skills/podcast-generation, .github/skills/podcast-generation and .opencode/skills/podcast-generation in your project.

What does Podcast Generation need to run?

Going by SKILL.md and its folder, Podcast Generation needs credentials named AZURE_OPENAI_AUDIO_API_KEY. Our summary lists: Python 3; A credential in AZURE_OPENAI_AUDIO_API_KEY.

Does Podcast Generation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Podcast Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Podcast Generation use?

Podcast Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Podcast Generation use?

About 913 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Podcast Generation?

Skills that share tags, products or a category with Podcast Generation: Azure Realtime Podcast Generation (microsoft/skills, 3.1k stars), Deepgram Python Text-to-Speech (deepgram/deepgram-python-sdk, 469 stars), Speech Engine (elevenlabs/skills, 482 stars) and Grok Realtime Voice Integration (cursor/plugins, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Podcast Generation?

aAAaqwq (a GitHub user) maintains it in aAAaqwq/AGI-Super-Team, which has 105 GitHub stars. The repository holds 167 skills in this directory. The repository was last updated on October 8, 2026.

Source: aAAaqwq/AGI-Super-Team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.