Agent skill

Podcast Generation

by bytedance in bytedance/deer-flow

Turns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script.

MITAuto-check passedMedia & Creative

Install Podcast Generation

skills CLI
$ npx skills add bytedance/deer-flow --skill podcast-generation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bytedance/deer-flow podcast-generation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public/podcast-generation .claude/skills/podcast-generation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
podcast-generation
GitHub stars
84k
Used in
2 other repos
Token cost
~2.2k tokens
SKILL.md length
741 words
Files
3 (incl. scripts)
Skills in repo
23
Repo updated
First seen
Licence
MIT

At a glance

Turns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script.

  • Works in 3 steps: Understand Requirements → Create Structured Script JSON → Execute Generation
  • Turning an article or report into a podcast episode
  • SKILL.md covers Overview, Core Capabilities, Workflow and Script JSON Format, plus 8 more sections
  • Runs Python scripts from its folder; calls python; reaches api.minimaxi.com; needs VOLCENGINE_TTS_ACCESS_TOKEN and MINIMAX_API_KEY

What it does

Given an article, report or other text, the skill has the agent write a structured JSON script of natural dialogue between a male and a female host, in English or Chinese, and then run the bundled generate.py once to synthesize the speech and mix the chunks into a final MP3. The script file goes in a workspace folder and the results in an outputs folder, and a tech-explainer template is bundled.

The script call takes the JSON script file, the output MP3 path and, recommended, a transcript file in Markdown, so you get a readable transcript beside the audio. The JSON has a locale of en or zh, an optional title used as the transcript heading and a list of lines. The agent should run it in one complete call without splitting the steps or reading the Python source, and the text-to-speech provider and its concurrency are picked automatically from environment variables.

When your agent uses it

  • Turning an article or report into a podcast episode
  • Producing a two-host conversation from documentation
  • Creating a Chinese-language podcast from written content
  • Generating a transcript alongside podcast audio

Example prompts

  • “Turn this quarterly report into a podcast episode and save it as review.mp3.”
  • “Make a two-host podcast from the README of this project, with a transcript.”
  • “把这篇文章做成一期播客,保存为 episode.mp3”

Requirements

  • Python, to run generate.py
  • A text-to-speech provider configured through environment variables

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Understand Requirements
  2. Create Structured Script JSON
  3. Execute Generation

What it can do on your machine

Read from SKILL.md and the folder at commit 8a3350a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.minimaxi.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VOLCENGINE_TTS_ACCESS_TOKEN
    • MINIMAX_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Podcast Generation loads about 2.2k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 741 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bytedance/deer-flow at commit 8a3350a, republished under its MIT licence (© bytedance). 741 words, ~2,156 tokens.

Download SKILL.mdSave it as .claude/skills/podcast-generation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
podcast-generation
description
Use this skill when the user requests to generate, create, or produce podcasts from text content. Converts written content into a two-host conversational podcast audio format with natural dialogue.

Podcast Generation Skill

Overview

This skill generates high-quality podcast audio from text content. The workflow includes creating a structured JSON script (conversational dialogue) and executing audio generation through text-to-speech synthesis.

Core Capabilities

  • Convert any text content (articles, reports, documentation) into podcast scripts
  • Generate natural two-host conversational dialogue (male and female hosts)
  • Synthesize speech audio using text-to-speech
  • Mix audio chunks into a final podcast MP3 file
  • Support both English and Chinese content

Workflow

Step 1: Understand Requirements

When a user requests podcast generation, identify:

  • Source content: The text/article/report to convert into a podcast
  • Language: English or Chinese (based on content)
  • Output location: Where to save the generated podcast
  • You don't need to check the folder under /mnt/user-data
Step 2: Create Structured Script JSON

Generate a structured JSON script file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}-script.json

The JSON structure:

json
{
  "locale": "en",
  "lines": [
    {"speaker": "male", "paragraph": "dialogue text"},
    {"speaker": "female", "paragraph": "dialogue text"}
  ]
}
Step 3: Execute Generation

Call the Python script:

bash
python /mnt/skills/public/podcast-generation/scripts/generate.py \
  --script-file /mnt/user-data/workspace/script-file.json \
  --output-file /mnt/user-data/outputs/generated-podcast.mp3 \
  --transcript-file /mnt/user-data/outputs/generated-podcast-transcript.md

Parameters:

  • --script-file: Absolute path to JSON script file (required)
  • --output-file: Absolute path to output MP3 file (required)
  • --transcript-file: Absolute path to output transcript markdown file (optional, but recommended)

[!IMPORTANT]

  • Execute the script in one complete call. Do NOT split the workflow into separate steps.
  • The script handles all TTS API calls and audio generation internally.
  • Do NOT read the Python file, just call it with the parameters.
  • Always include --transcript-file to generate a readable transcript for the user.
  • The TTS provider and its concurrency are selected automatically from environment variables — you do not choose or tune them.

Script JSON Format

The script JSON file must follow this structure:

json
{
  "title": "The History of Artificial Intelligence",
  "locale": "en",
  "lines": [
    {"speaker": "male", "paragraph": "Hello Deer! Welcome back to another episode."},
    {"speaker": "female", "paragraph": "Hey everyone! Today we have an exciting topic to discuss."},
    {"speaker": "male", "paragraph": "That's right! We're going to talk about..."}
  ]
}

Fields:

  • title: Title of the podcast episode (optional, used as heading in transcript)
  • locale: Language code - "en" for English or "zh" for Chinese
  • lines: Array of dialogue lines
    • speaker: Either "male" or "female"
    • paragraph: The dialogue text for this speaker

Script Writing Guidelines

When creating the script JSON, follow these guidelines:

Format Requirements
  • Only two hosts: male and female, alternating naturally
  • Target runtime: approximately 10 minutes of dialogue (around 40-60 lines)
  • Start with the male host saying a greeting that includes "Hello Deer"
Tone & Style
  • Natural, conversational dialogue - like two friends chatting
  • Use casual expressions and conversational transitions
  • Avoid overly formal language or academic tone
  • Include reactions, follow-up questions, and natural interjections
Content Guidelines
  • Frequent back-and-forth between hosts
  • Keep sentences short and easy to follow when spoken
  • Plain text only - no markdown formatting in the output
  • Translate technical concepts into accessible language
  • No mathematical formulas, code, or complex notation
  • Make content engaging and accessible for audio-only listeners
  • Exclude meta information like dates, author names, or document structure

Podcast Generation Example

User request: "Generate a podcast about the history of artificial intelligence"

Step 1: Create script file /mnt/user-data/workspace/ai-history-script.json:

json
{
  "title": "The History of Artificial Intelligence",
  "locale": "en",
  "lines": [
    {"speaker": "male", "paragraph": "Hello Deer! Welcome back to another fascinating episode. Today we're diving into something that's literally shaping our future - the history of artificial intelligence."},
    {"speaker": "female", "paragraph": "Oh, I love this topic! You know, AI feels so modern, but it actually has roots going back over seventy years."},
    {"speaker": "male", "paragraph": "Exactly! It all started back in the 1950s. The term artificial intelligence was actually coined by John McCarthy in 1956 at a famous conference at Dartmouth."},
    {"speaker": "female", "paragraph": "Wait, so they were already thinking about machines that could think back then? That's incredible!"},
    {"speaker": "male", "paragraph": "Right? The early pioneers were so optimistic. They thought we'd have human-level AI within a generation."},
    {"speaker": "female", "paragraph": "But things didn't quite work out that way, did they?"},
    {"speaker": "male", "paragraph": "No, not at all. The 1970s brought what's called the first AI winter..."}
  ]
}

Step 2: Execute generation:

bash
python /mnt/skills/public/podcast-generation/scripts/generate.py \
  --script-file /mnt/user-data/workspace/ai-history-script.json \
  --output-file /mnt/user-data/outputs/ai-history-podcast.mp3 \
  --transcript-file /mnt/user-data/outputs/ai-history-transcript.md

This will generate:

  • ai-history-podcast.mp3: The audio podcast file
  • ai-history-transcript.md: A readable markdown transcript of the podcast
Show full SKILL.md (281 more words)Show less

Specific Templates

Read the following template file only when matching the user request.

  • Tech Explainer - For converting technical documentation and tutorials

Output Format

The generated podcast follows the "Hello Deer" format:

  • Two hosts: one male, one female
  • Natural conversational dialogue
  • Starts with "Hello Deer" greeting
  • Target duration: approximately 10 minutes
  • Alternating speakers for engaging flow

Output Handling

After generation:

  • Podcasts and transcripts are saved in /mnt/user-data/outputs/
  • Share both the podcast MP3 and transcript MD with user using present_files tool
  • Provide brief description of the generation result (topic, duration, hosts)
  • Offer to regenerate if adjustments needed

Requirements

The following environment variables must be set:

  • For Volcengine: VOLCENGINE_TTS_APPID and VOLCENGINE_TTS_ACCESS_TOKEN
  • For MiniMax: MINIMAX_API_KEY
  • VOLCENGINE_TTS_CLUSTER: Volcengine TTS cluster (optional, defaults to "volcano_tts")
  • VOLCENGINE_TTS_VOICE_TYPE_MALE: Volcengine male voice type (optional, defaults to zh_male_yangguangqingnian_moon_bigtts)
  • VOLCENGINE_TTS_VOICE_TYPE_FEMALE: Volcengine female voice type (optional, defaults to zh_female_sajiaonvyou_moon_bigtts)

Voice type overrides are trimmed; unset or blank values use the listed defaults.

Notes

  • Always execute the full pipeline in one call - no need to test individual steps or worry about timeouts
  • The script JSON should match the content language (en or zh)
  • Technical content should be simplified for audio accessibility in the script
  • Complex notations (formulas, code) should be translated to plain language in the script
  • Long content may result in longer podcasts

Providers (Volcengine / MiniMax)

Auto-selected by environment variables:

  • VOLCENGINE_TTS_APPID + VOLCENGINE_TTS_ACCESS_TOKEN set → Volcengine TTS (default).
  • Only MINIMAX_API_KEY set → MiniMax TTS (/v1/t2a_v2).
  • Force with PODCAST_GENERATION_PROVIDER=volcengine|minimax.

MiniMax overrides: MINIMAX_API_HOST (default https://api.minimaxi.com), MINIMAX_TTS_MODEL (default speech-2.6-hd), MINIMAX_TTS_VOICE_MALE (default male-qn-qingse), MINIMAX_TTS_VOICE_FEMALE (default female-tianmei).

Concurrency is owned by each provider internally — MiniMax runs single-threaded to reduce rate-limit failures, Volcengine uses 4 workers. There is no caller-facing concurrency knob; transient rate limits are handled by automatic retry with backoff.

© bytedance, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/public/podcast-generation of bytedance/deer-flow.

  • SKILL.md
  • scripts/generate.py
  • templates/tech-explainer.md

Open the folder on GitHubat commit 8a3350a

Used in 2 other repositories

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in bytedance/deer-flow, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Podcast Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Podcast Generation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Podcast Generation this skillbytedance/deer-flow84k2 repos~2.2kAutomated safety check: PassMIT
Video Podcast Makerdtsola/xiaoyaosearch1k—~3.4kAutomated safety check: PassMIT
Azure Realtime Podcast Generationmicrosoft/skills3.1k1 repos~947Automated safety check: PassMIT
Podcast Pipelineericosiu/ai-marketing-skills3.6k1 repos~2.6kAutomated safety check: PassMIT
Audio Script WriterLeoYeAI/openclaw-master-skills2.2k—~4.1kAutomated safety check: NotesMIT
Youtube Downloaderguia-matthieu/clawfu-skills150—~1.3kAutomated safety check: PassMIT

Similar skills

  • Video Podcast Maker

    dtsola/xiaoyaosearch

    Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

    1k GitHub stars~3.4k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Media & CreativeAuto-check passed
  • Podcast Pipeline

    ericosiu/ai-marketing-skills

    Podcast-to-Everything content pipeline. An agent skill from ericosiu/ai-marketing-skills.

    3.6k GitHub starsUsed in 1 repo~2.6k tokens
    Writing & ContentAuto-check passed
  • Audio Script Writer

    LeoYeAI/openclaw-master-skills

    Convert written medical content into podcast or video scripts optimized for audio delivery.

    2.2k GitHub stars~4.1k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Youtube Downloader

    guia-matthieu/clawfu-skills

    Download and process YouTube content for research. An agent skill from guia-matthieu/clawfu-skills.

    150 GitHub stars~1.3k tokensUpdated 10 days ago
    Writing & ContentAuto-check passed
  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings

More from bytedance/deer-flow

All 23 skills in this repo
  • Vercel Deploy

    bytedance/deer-flow

    Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.

    84k GitHub starsUsed in 10 repos~797 tokens
    Auto-check passed
  • GitHub Deep Research

    bytedance/deer-flow

    Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.

    84k GitHub starsUsed in 4 repos~1.3k tokens
    Auto-check passed
  • Chart Visualization

    bytedance/deer-flow

    Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.

    84k GitHub starsUsed in 1 repo~840 tokens
    Auto-check passed
  • Excel and CSV Data Analysis

    bytedance/deer-flow

    Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.

    84k GitHub starsUsed in 4 repos~2.2k tokens
    Auto-check passed
  • Structured Image Generation

    bytedance/deer-flow

    Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.

    84k GitHub starsUsed in 4 repos~2.9k tokens
    Auto-check passed
  • DeerFlow Smoke Test

    bytedance/deer-flow

    Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.

    84k GitHub stars~2.5k tokensUpdated today
    Auto-check: notes

Works with

Questions about Podcast Generation

What does Podcast Generation do?

Turns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script. py once to synthesize the speech and mix the chunks into a final MP3. The script file goes in a workspace folder and the results in an outputs folder, and a tech-explainer template is bundled.

When should I use Podcast Generation?

Podcast Generation fits situations like: turning an article or report into a podcast episode; producing a two-host conversation from documentation; creating a Chinese-language podcast from written content; generating a transcript alongside podcast audio.

How do I install Podcast Generation in Claude Code?

Run `npx skills add bytedance/deer-flow --skill podcast-generation -a claude-code`. Or copy the skill folder (skills/public/podcast-generation in bytedance/deer-flow) into .claude/skills/podcast-generation in your project. Claude Code loads it when a task matches its description.

How do I install Podcast Generation in Codex?

Run `npx skills add bytedance/deer-flow --skill podcast-generation -a codex`. Or copy the skill folder (skills/public/podcast-generation in bytedance/deer-flow) into .agents/skills/podcast-generation in your project. Codex loads it when a task matches its description.

Can I use Podcast Generation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bytedance/deer-flow --skill podcast-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/podcast-generation, .gemini/skills/podcast-generation, .github/skills/podcast-generation and .opencode/skills/podcast-generation in your project.

What does Podcast Generation need to run?

Going by SKILL.md and its folder, Podcast Generation needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named VOLCENGINE_TTS_ACCESS_TOKEN and MINIMAX_API_KEY. Our summary lists: Python, to run generate.py; A text-to-speech provider configured through environment variables.

Does Podcast Generation access the network?

SKILL.md names 1 domain. In commands or code: api.minimaxi.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Podcast Generation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Podcast Generation use?

Podcast Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Podcast Generation use?

About 2.2k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Podcast Generation?

Skills that share tags, products or a category with Podcast Generation: Video Podcast Maker (dtsola/xiaoyaosearch, 1k stars), Azure Realtime Podcast Generation (microsoft/skills, 3.1k stars), Podcast Pipeline (ericosiu/ai-marketing-skills, 3.6k stars) and Audio Script Writer (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Podcast Generation?

bytedance (a GitHub organization) maintains it in bytedance/deer-flow, which has 83,674 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 11, 2026.

Source: bytedance/deer-flow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.