Video Podcast Maker
dtsola/xiaoyaosearch
Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.
Turns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script.
$ npx skills add bytedance/deer-flow --skill podcast-generation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install bytedance/deer-flow podcast-generation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/public/podcast-generation .claude/skills/podcast-generation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "podcast-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/podcast-generation into .claude/skills/podcast-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "podcast-generation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/bytedance/deer-flow/tree/main/skills/public/podcast-generationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add bytedance/deer-flow --skill podcast-generation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install bytedance/deer-flow podcast-generation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/public/podcast-generation .agents/skills/podcast-generation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "podcast-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/podcast-generation into .agents/skills/podcast-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "podcast-generation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bytedance/deer-flow --skill podcast-generation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install bytedance/deer-flow podcast-generation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/public/podcast-generation .cursor/skills/podcast-generation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "podcast-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/podcast-generation into .cursor/skills/podcast-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "podcast-generation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/bytedance/deer-flow.git --path skills/public/podcast-generation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add bytedance/deer-flow --skill podcast-generation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install bytedance/deer-flow podcast-generation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/public/podcast-generation .gemini/skills/podcast-generation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "podcast-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/podcast-generation into .gemini/skills/podcast-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "podcast-generation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install bytedance/deer-flow podcast-generationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add bytedance/deer-flow --skill podcast-generation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/public/podcast-generation .github/skills/podcast-generation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "podcast-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/podcast-generation into .github/skills/podcast-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "podcast-generation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add bytedance/deer-flow --skill podcast-generation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install bytedance/deer-flow podcast-generation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/bytedance/deer-flow.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/public/podcast-generation .opencode/skills/podcast-generation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "podcast-generation" agent skill from https://github.com/bytedance/deer-flow/tree/main/skills/public/podcast-generation into .opencode/skills/podcast-generation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "podcast-generation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
podcast-generationTurns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script.
Given an article, report or other text, the skill has the agent write a structured JSON script of natural dialogue between a male and a female host, in English or Chinese, and then run the bundled generate.py once to synthesize the speech and mix the chunks into a final MP3. The script file goes in a workspace folder and the results in an outputs folder, and a tech-explainer template is bundled.
The script call takes the JSON script file, the output MP3 path and, recommended, a transcript file in Markdown, so you get a readable transcript beside the audio. The JSON has a locale of en or zh, an optional title used as the transcript heading and a list of lines. The agent should run it in one complete call without splitting the steps or reading the Python source, and the text-to-speech provider and its concurrency are picked automatically from environment variables.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8a3350a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.minimaxi.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
VOLCENGINE_TTS_ACCESS_TOKENMINIMAX_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Podcast Generation loads about 2.2k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 741 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from bytedance/deer-flow at commit 8a3350a, republished under its MIT licence (© bytedance). 741 words, ~2,156 tokens.
.claude/skills/podcast-generation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill generates high-quality podcast audio from text content. The workflow includes creating a structured JSON script (conversational dialogue) and executing audio generation through text-to-speech synthesis.
When a user requests podcast generation, identify:
/mnt/user-dataGenerate a structured JSON script file in /mnt/user-data/workspace/ with naming pattern: {descriptive-name}-script.json
The JSON structure:
{
"locale": "en",
"lines": [
{"speaker": "male", "paragraph": "dialogue text"},
{"speaker": "female", "paragraph": "dialogue text"}
]
}Call the Python script:
python /mnt/skills/public/podcast-generation/scripts/generate.py \
--script-file /mnt/user-data/workspace/script-file.json \
--output-file /mnt/user-data/outputs/generated-podcast.mp3 \
--transcript-file /mnt/user-data/outputs/generated-podcast-transcript.mdParameters:
--script-file: Absolute path to JSON script file (required)--output-file: Absolute path to output MP3 file (required)--transcript-file: Absolute path to output transcript markdown file (optional, but recommended)[!IMPORTANT]
- Execute the script in one complete call. Do NOT split the workflow into separate steps.
- The script handles all TTS API calls and audio generation internally.
- Do NOT read the Python file, just call it with the parameters.
- Always include
--transcript-fileto generate a readable transcript for the user.- The TTS provider and its concurrency are selected automatically from environment variables — you do not choose or tune them.
The script JSON file must follow this structure:
{
"title": "The History of Artificial Intelligence",
"locale": "en",
"lines": [
{"speaker": "male", "paragraph": "Hello Deer! Welcome back to another episode."},
{"speaker": "female", "paragraph": "Hey everyone! Today we have an exciting topic to discuss."},
{"speaker": "male", "paragraph": "That's right! We're going to talk about..."}
]
}Fields:
title: Title of the podcast episode (optional, used as heading in transcript)locale: Language code - "en" for English or "zh" for Chineselines: Array of dialogue linesspeaker: Either "male" or "female"paragraph: The dialogue text for this speakerWhen creating the script JSON, follow these guidelines:
User request: "Generate a podcast about the history of artificial intelligence"
Step 1: Create script file /mnt/user-data/workspace/ai-history-script.json:
{
"title": "The History of Artificial Intelligence",
"locale": "en",
"lines": [
{"speaker": "male", "paragraph": "Hello Deer! Welcome back to another fascinating episode. Today we're diving into something that's literally shaping our future - the history of artificial intelligence."},
{"speaker": "female", "paragraph": "Oh, I love this topic! You know, AI feels so modern, but it actually has roots going back over seventy years."},
{"speaker": "male", "paragraph": "Exactly! It all started back in the 1950s. The term artificial intelligence was actually coined by John McCarthy in 1956 at a famous conference at Dartmouth."},
{"speaker": "female", "paragraph": "Wait, so they were already thinking about machines that could think back then? That's incredible!"},
{"speaker": "male", "paragraph": "Right? The early pioneers were so optimistic. They thought we'd have human-level AI within a generation."},
{"speaker": "female", "paragraph": "But things didn't quite work out that way, did they?"},
{"speaker": "male", "paragraph": "No, not at all. The 1970s brought what's called the first AI winter..."}
]
}Step 2: Execute generation:
python /mnt/skills/public/podcast-generation/scripts/generate.py \
--script-file /mnt/user-data/workspace/ai-history-script.json \
--output-file /mnt/user-data/outputs/ai-history-podcast.mp3 \
--transcript-file /mnt/user-data/outputs/ai-history-transcript.mdThis will generate:
ai-history-podcast.mp3: The audio podcast fileai-history-transcript.md: A readable markdown transcript of the podcastRead the following template file only when matching the user request.
The generated podcast follows the "Hello Deer" format:
After generation:
/mnt/user-data/outputs/present_files toolThe following environment variables must be set:
VOLCENGINE_TTS_APPID and VOLCENGINE_TTS_ACCESS_TOKENMINIMAX_API_KEYVOLCENGINE_TTS_CLUSTER: Volcengine TTS cluster (optional, defaults to "volcano_tts")VOLCENGINE_TTS_VOICE_TYPE_MALE: Volcengine male voice type (optional, defaults to zh_male_yangguangqingnian_moon_bigtts)VOLCENGINE_TTS_VOICE_TYPE_FEMALE: Volcengine female voice type (optional, defaults to zh_female_sajiaonvyou_moon_bigtts)Voice type overrides are trimmed; unset or blank values use the listed defaults.
Auto-selected by environment variables:
VOLCENGINE_TTS_APPID + VOLCENGINE_TTS_ACCESS_TOKEN set → Volcengine TTS (default).MINIMAX_API_KEY set → MiniMax TTS (/v1/t2a_v2).PODCAST_GENERATION_PROVIDER=volcengine|minimax.MiniMax overrides: MINIMAX_API_HOST (default https://api.minimaxi.com),
MINIMAX_TTS_MODEL (default speech-2.6-hd), MINIMAX_TTS_VOICE_MALE
(default male-qn-qingse), MINIMAX_TTS_VOICE_FEMALE (default female-tianmei).
Concurrency is owned by each provider internally — MiniMax runs single-threaded to reduce rate-limit failures, Volcengine uses 4 workers. There is no caller-facing concurrency knob; transient rate limits are handled by automatic retry with backoff.
© bytedance, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (scripts) in skills/public/podcast-generation of bytedance/deer-flow.
Open the folder on GitHubat commit 8a3350a
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in bytedance/deer-flow, which our catalogue first saw on October 7, 2026.
Podcast Generation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Podcast Generation this skillbytedance/deer-flow | 84k | 2 repos | ~2.2k | Automated safety check: Pass | MIT | |
| Video Podcast Makerdtsola/xiaoyaosearch | 1k | — | ~3.4k | Automated safety check: Pass | MIT | |
| Azure Realtime Podcast Generationmicrosoft/skills | 3.1k | 1 repos | ~947 | Automated safety check: Pass | MIT | |
| Podcast Pipelineericosiu/ai-marketing-skills | 3.6k | 1 repos | ~2.6k | Automated safety check: Pass | MIT | |
| Audio Script WriterLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.1k | Automated safety check: Notes | MIT | |
| Youtube Downloaderguia-matthieu/clawfu-skills | 150 | — | ~1.3k | Automated safety check: Pass | MIT |
dtsola/xiaoyaosearch
Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.
microsoft/skills
Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.
ericosiu/ai-marketing-skills
Podcast-to-Everything content pipeline. An agent skill from ericosiu/ai-marketing-skills.
LeoYeAI/openclaw-master-skills
Convert written medical content into podcast or video scripts optimized for audio delivery.
guia-matthieu/clawfu-skills
Download and process YouTube content for research. An agent skill from guia-matthieu/clawfu-skills.
harry0703/MoneyPrinterTurbo
Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.
bytedance/deer-flow
Deploys a project to Vercel with one script and no login, then returns a live preview URL and a claim link for moving the deployment into your own Vercel account.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
bytedance/deer-flow
Picks a suitable chart type from 26 options for your data, maps the data to that chart's parameters and generates a chart image through a JavaScript script.
bytedance/deer-flow
Analyzes uploaded Excel and CSV files with SQL through DuckDB, producing schema inspections, statistical summaries and exports to CSV, JSON or Markdown.
bytedance/deer-flow
Turns an image request into a structured JSON prompt and runs a bundled Python script to generate the picture, optionally guided by reference images.
bytedance/deer-flow
Walks through an end-to-end smoke test of a DeerFlow deployment: pull the latest code, deploy with Docker or locally, verify services, run health checks and write a report.
Works with
Categories
Turns written content into a two-host conversational podcast MP3 with a transcript, by drafting a JSON script and running a text-to-speech script. py once to synthesize the speech and mix the chunks into a final MP3. The script file goes in a workspace folder and the results in an outputs folder, and a tech-explainer template is bundled.
Podcast Generation fits situations like: turning an article or report into a podcast episode; producing a two-host conversation from documentation; creating a Chinese-language podcast from written content; generating a transcript alongside podcast audio.
Run `npx skills add bytedance/deer-flow --skill podcast-generation -a claude-code`. Or copy the skill folder (skills/public/podcast-generation in bytedance/deer-flow) into .claude/skills/podcast-generation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add bytedance/deer-flow --skill podcast-generation -a codex`. Or copy the skill folder (skills/public/podcast-generation in bytedance/deer-flow) into .agents/skills/podcast-generation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bytedance/deer-flow --skill podcast-generation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/podcast-generation, .gemini/skills/podcast-generation, .github/skills/podcast-generation and .opencode/skills/podcast-generation in your project.
Going by SKILL.md and its folder, Podcast Generation needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named VOLCENGINE_TTS_ACCESS_TOKEN and MINIMAX_API_KEY. Our summary lists: Python, to run generate.py; A text-to-speech provider configured through environment variables.
SKILL.md names 1 domain. In commands or code: api.minimaxi.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Podcast Generation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Podcast Generation: Video Podcast Maker (dtsola/xiaoyaosearch, 1k stars), Azure Realtime Podcast Generation (microsoft/skills, 3.1k stars), Podcast Pipeline (ericosiu/ai-marketing-skills, 3.6k stars) and Audio Script Writer (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
bytedance (a GitHub organization) maintains it in bytedance/deer-flow, which has 83,674 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 11, 2026.
Source: bytedance/deer-flow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.