Agent skill

Video Podcast Maker

by dtsola in dtsola/xiaoyaosearch

Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

MITAuto-check passedMedia & Creative

Install Video Podcast Maker

skills CLI
$ npx skills add dtsola/xiaoyaosearch --skill video-podcast-maker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dtsola/xiaoyaosearch video-podcast-maker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dtsola/xiaoyaosearch.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/video-podcast-maker .claude/skills/video-podcast-maker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-podcast-maker
GitHub stars
1k
Token cost
~3.4k tokens
SKILL.md length
979 words
Files
90 (incl. scripts, references, assets)
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

  • Works in 2 steps: Step 9: Launch Remotion Studio — user… → Step 10: Only triggered when user…
  • Making a narrated knowledge video from a single topic
  • SKILL.md covers Quick Start, Design Learning (Optional), Auto Update Check and Prerequisites Check, plus 9 more sections
  • Calls git, npx and python3

What it does

Given a topic, the skill runs an automated pipeline from research and script through TTS narration and Remotion composition to an MP4 with background music, aimed at horizontal knowledge videos for Bilibili in either 4K or 1080p with a clean white style by default. The tech stack is the coding agent, a TTS backend, Remotion and FFmpeg, and it depends on a separate Remotion best-practices skill being loaded first.

Optional design learning extracts visual patterns from a reference video or image and saves them as style profiles that can be listed, applied or deleted. A prerequisites script checks the setup, resolving the TTS backend from an environment variable, then a preferences file, then the edge default, and validating only what that backend needs. The skill also checks for updates at most once a day and asks before pulling. Assets include sample music, intro clips, a workflow diagram and design and troubleshooting references.

When your agent uses it

  • Making a narrated knowledge video from a single topic
  • Learning visual design patterns from a reference video or image
  • Producing a Bilibili-style horizontal video with background music
  • Saving or reusing a video style profile

Example prompts

  • “Make a video podcast about how vector databases work.”
  • “Learn the visual style of this reference video and apply it to my next episode.”
  • “Create a video podcast on AI agent tutorials with a clean white style.”

Requirements

  • Remotion with the `remotion-best-practices` skill
  • FFmpeg
  • A TTS backend, with edge as the default
  • Python 3 for the prerequisites check

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Step 9: Launch Remotion Studio — user reviews in real-time, requests changes until satisfied
  2. Step 10: Only triggered when user explicitly says "render 4K" / "render final version"

What it can do on your machine

Read from SKILL.md and the folder at commit 0a9c6fc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • git
    • npx
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Podcast Maker loads about 3.4k tokens when it runs, and up to ~21k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 979 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~21k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from dtsola/xiaoyaosearch at commit 0a9c6fc, republished under its MIT licence (© dtsola). 979 words, ~3,389 tokens.

Download SKILL.mdSave it as .claude/skills/video-podcast-maker/SKILL.md (or your agent's skills folder). This skill also uses 89 other files; get the full folder from GitHub.
name
video-podcast-maker
description
Use when the user gives a topic and wants an automated video podcast created, or asks to learn visual design patterns from a reference video/image. Produces 4K video via research → script → TTS → Remotion → MP4 + BGM.
argument-hint
[topic]
effort
high
author
Agents365-ai
category
Content Creation
version
2.0.0
created
2025-01-27
updated
2026-04-03
bilibili
https://space.bilibili.com/441831884
github
https://github.com/Agents365-ai/video-podcast-maker
dependencies
remotion-best-practices

REQUIRED: Load Remotion Best Practices First

This skill depends on remotion-best-practices. You MUST invoke it before proceeding:

Invoke the skill/tool named: remotion-best-practices

Video Podcast Maker

Quick Start

Open your coding agent and say: "Make a video podcast about $ARGUMENTS"

Or invoke directly: /video-podcast-maker AI Agent tutorial


Design Learning (Optional)

Extract visual design patterns from reference videos or images and apply them to new video compositions. Skip this section unless the user provides a reference video/image or asks to save/list/delete style profiles.

→ See references/design-learning.md for commands, reference-library management, style-profile management, and integration with Pre-workflow / Step 9.


Auto Update Check

Agent behavior: Check for updates at most once per day (throttled by timestamp file). Before any shell command that reads files from this skill, resolve SKILL_DIR to the directory containing SKILL.md. If your agent exposes a built-in skill directory variable such as ${CLAUDE_SKILL_DIR}, you may map it to SKILL_DIR.

bash
SKILL_DIR="${SKILL_DIR:-${CLAUDE_SKILL_DIR}}"
STAMP="${SKILL_DIR}/.last_update_check"
NOW=$(date +%s)
LAST=$(cat "$STAMP" 2>/dev/null || echo 0)
if [ ! -d "${SKILL_DIR}/.git" ]; then
  echo "MANUAL_INSTALL"
elif [ $((NOW - LAST)) -gt 86400 ]; then
  timeout 5 git -C "${SKILL_DIR}" fetch --quiet 2>/dev/null || true
  LOCAL=$(git -C "${SKILL_DIR}" rev-parse HEAD 2>/dev/null)
  REMOTE=$(git -C "${SKILL_DIR}" rev-parse origin/main 2>/dev/null)
  echo "$NOW" > "$STAMP"
  if [ -n "$LOCAL" ] && [ -n "$REMOTE" ] && [ "$LOCAL" != "$REMOTE" ]; then
    echo "UPDATE_AVAILABLE"
  else
    echo "UP_TO_DATE"
  fi
else
  echo "SKIPPED_RECENT_CHECK"
fi
  • Update available: Ask the user whether to pull updates. Yes → git -C "${SKILL_DIR}" pull. No → continue.
  • Up to date / Skipped: Continue silently.
  • Manual install (no .git directory — skill was installed via tarball/zip/cp): Continue silently. Auto-update is disabled; the user must reinstall manually to update.

Prerequisites Check

!python3 "${SKILL_DIR}/scripts/check_prereqs.py"

If MISSING reported above, see README.md for full setup instructions (install commands, API key setup, Remotion project init). The check is backend-aware: backend is resolved as TTS_BACKEND env var → user_prefs.json (global.tts.backend) → edge default, then only env vars required by that backend are validated.


Overview

Automated pipeline for professional Bilibili horizontal knowledge videos from a topic.

Target: Bilibili horizontal video (16:9)

  • Resolution: 3840×2160 (4K) or 1920×1080 (1080p)
  • Style: Clean white (default)

Tech stack: Coding agent + TTS backend + Remotion + FFmpeg

Output Specs
ParameterHorizontal (16:9)Vertical (9:16)
Resolution3840×2160 (4K)2160×3840 (4K)
Frame rate30 fps30 fps
EncodingH.264, 16MbpsH.264, 16Mbps
AudioAAC, 192kbpsAAC, 192kbps
Duration1-15 min60-90s (highlight)

Execution Modes

Agent behavior: Detect user intent at workflow start:

  • "Make a video about..." / no special instructions → Auto Mode
  • "I want to control each step" / mentions interactive → Interactive Mode
  • Default: Auto Mode
Auto Mode (Default)

Full pipeline with sensible defaults. Mandatory stop at Step 9:

  1. Step 9: Launch Remotion Studio — user reviews in real-time, requests changes until satisfied
  2. Step 10: Only triggered when user explicitly says "render 4K" / "render final version"
StepDecisionAuto Default
3Title positiontop-center
5Media assetsSkip (text-only animations)
7Thumbnail methodRemotion-generated (16:9 + 4:3)
9Outro animationPre-made MP4 (white/black by theme)
9Preview methodRemotion Studio (mandatory)
12SubtitlesSkip
14CleanupAuto-clean temp files

Users can override any default in their initial request:

  • "make a video about AI, burn subtitles" → auto + subtitles on
  • "use dark theme, AI thumbnails" → auto + dark + imagen
  • "need screenshots" → auto + media collection enabled
Interactive Mode

Prompts at each decision point. Activated by:

  • "interactive mode" / "I want to choose each option"
  • User explicitly requests control

Technical Rules

Hard constraints for video production. Visual design remains the agent's creative freedom within these rules:

RuleRequirement
Single ProjectAll videos under videos/{name}/ in user's Remotion project. NEVER create a new project per video.
4K Output3840×2160, use scale(2) wrapper over 1920×1080 design space
Content Width≥85% of screen width
Bottom Safe ZoneBottom 100px reserved for subtitles
Audio SyncAll animations driven by timing.json timestamps
ThumbnailMUST generate 16:9 (1920×1080) AND 4:3 (1200×900). Centered layout, title ≥120px, icons ≥120px, fill most of canvas. See design-guide.md.
FontPingFang SC / Noto Sans SC for Chinese text
Studio Before RenderMUST launch remotion studio for user review. NEVER render 4K until user explicitly confirms ("render 4K", "render final").

Show full SKILL.md (386 more words)Show less

Additional Resources

Load these files on demand — do NOT load all at once:

  • references/workflow-steps.md: Index of the 15-step workflow split across three phase files. Load at workflow start to locate which phase file to pull:
  • references/design-guide.md: Visual minimums, typography, layout patterns, checklists. MUST load before Step 9.
  • references/design-learning.md: Extracting visual patterns from reference videos/images, style profiles. Load only when the user provides a reference or manages profiles.
  • references/troubleshooting.md: Error fixes, BGM options, preference commands, preference learning. Load on error or user request.
  • examples/: Real production video projects. The agent may reference these for composition structure and timing.json format.

Directory Structure

project-root/                           # Remotion project root
├── src/remotion/                       # Remotion source
│   ├── compositions/                   # Video composition definitions
│   ├── Root.tsx                        # Remotion entry
│   └── index.ts                        # Exports
│
├── public/                             # Remotion default (unused — use --public-dir videos/{name}/)
│
├── videos/{video-name}/                # Video project assets
│   ├── topic_definition.md             # Step 1
│   ├── topic_research.md               # Step 2
│   ├── podcast.txt                     # Step 4: narration script
│   ├── podcast_audio.wav               # Step 8: TTS audio
│   ├── podcast_audio.srt               # Step 8: subtitles
│   ├── timing.json                     # Step 8: timeline
│   ├── thumbnail_*.png                 # Step 7
│   ├── output.mp4                      # Step 10
│   ├── video_with_bgm.mp4             # Step 11
│   ├── final_video.mp4                 # Step 12: final output
│   └── bgm.mp3                         # Background music
│
└── remotion.config.ts

Important: Always use --public-dir and full output path for Remotion render:

bash
npx remotion render src/remotion/index.ts CompositionId videos/{name}/output.mp4 --public-dir videos/{name}/
Naming Rules

Video name {video-name}: lowercase English, hyphen-separated (e.g., reference-manager-comparison)

Section name {section}: lowercase English, underscore-separated, matches [SECTION:xxx]

Thumbnail naming (16:9 AND 4:3 both required):

Type16:94:3
Remotionthumbnail_remotion_16x9.pngthumbnail_remotion_4x3.png
AIthumbnail_ai_16x9.pngthumbnail_ai_4x3.png
Public Directory

Use --public-dir videos/{name}/ for all Remotion commands. Each video's assets (timing.json, podcast_audio.wav, bgm.mp3) stay in its own directory — no copying to public/ needed. This enables parallel renders of different videos.

bash
# All render/studio/still commands use --public-dir
npx remotion studio src/remotion/index.ts --public-dir videos/{name}/
npx remotion render src/remotion/index.ts CompositionId videos/{name}/output.mp4 --public-dir videos/{name}/ --video-bitrate 16M
npx remotion still src/remotion/index.ts Thumbnail16x9 videos/{name}/thumbnail.png --public-dir videos/{name}/

Workflow

Progress Tracking

At Step 1 start, use your agent's task tracker (Claude Code TaskCreate / Codex todo list / equivalent) to create one task per step. Mark in_progress on start, completed on finish. Files in videos/{name}/ (e.g. podcast.txt, timing.json, output.mp4) act as the durable record of what completed — if a session is interrupted, inspect the directory to determine where to resume.

 1. Define topic direction → topic_definition.md
 2. Research topic → topic_research.md
 3. Design video sections (5-7 chapters)
 4. Write narration script → podcast.txt
 4.5. Pronunciation pre-flight (zh-CN only) → videos/{name}/phonemes.json
 5. Collect media assets → media_manifest.json
 6. Generate publish info (Part 1) → publish_info.md
 7. Generate thumbnails (16:9 + 4:3) → thumbnail_*.png
 8. Generate TTS audio → podcast_audio.wav, timing.json
 9. Create Remotion composition + Studio preview (mandatory stop)
10. Render 4K video (only on user request) → output.mp4
11. Mix background music → video_with_bgm.mp4
12. Add subtitles (optional) → final_video.mp4
13. Complete publish info (Part 2) → chapter timestamps
14. Verify output & cleanup
15. Generate vertical shorts (optional) → shorts/
Validation Checkpoints

After Step 8 (TTS):

  • podcast_audio.wav exists and plays correctly
  • timing.json has all sections with correct timestamps
  • podcast_audio.srt encoding is UTF-8

After Step 10 (Render):

  • output.mp4 resolution is 3840x2160
  • Audio-video sync verified
  • No black frames

Key Commands Reference

See CLAUDE.md for the full command reference (TTS, Remotion, FFmpeg, shorts generation).


User Preference System

Skill learns and applies preferences automatically. See references/troubleshooting.md for commands and learning details.

Storage Files
FilePurpose
user_prefs.jsonLearned preferences (auto-created from template)
user_prefs.template.jsonDefault values
prefs_schema.jsonJSON schema definition
Priority
Final = merge(Root.tsx defaults < global < topic_patterns[type] < current instructions)
User Commands
CommandEffect
"show preferences"Show current preferences
"reset preferences"Reset to defaults
"save as X default"Save to topic_patterns

Troubleshooting & Preferences

Full reference: Read references/troubleshooting.md on errors, preference questions, or BGM options.

© dtsola, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 89 other files (scripts, references, assets) in .claude/skills/video-podcast-maker of dtsola/xiaoyaosearch.

  • SKILL.md
  • .env.example
  • .gitignore
  • .python-version
  • LICENSE
  • README.md
  • README_CN.md
  • assets/bilibili-triple-black.mp4
  • assets/bilibili-triple-white.mp4
  • assets/perfect-beauty-191271.mp3
  • assets/snow-stevekaldes-piano-397491.mp3
  • assets/workflow.mmd
  • assets/workflow.png
  • package.json
  • phonemes.template.json
  • prefs_schema.json
  • references/design-guide.md
  • references/design-learning.md
  • references/troubleshooting.md
  • … and 71 more

Open the folder on GitHubat commit 0a9c6fc

Compare with similar skills

Video Podcast Maker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Podcast Maker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Podcast Maker this skilldtsola/xiaoyaosearch1k—~3.4kAutomated safety check: PassMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT
Whiteboard Videognipbao/codex-whiteboard-video-skill327—~7.2kAutomated safety check: NotesMIT
Super Video MakerBomx/super-video-maker-skill310—~11kAutomated safety check: NotesNone
Video Podcast Maker NanoAgents365-ai/video-podcast-maker1.7k—~3.6kAutomated safety check: PassMIT
Muapi DirectorAnil-matcha/vox-ai-motion-graphics-generator246—~679Automated safety check: PassNone

Similar skills

  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • Whiteboard Video

    gnipbao/codex-whiteboard-video-skill

    Generate smooth hand-drawn whiteboard and story videos directly inside Codex from text scripts, GPT Image 2 color storyboards, scene plans, SVGs, line art, or local images.

    327 GitHub stars~7.2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Video Podcast Maker Nano

    Agents365-ai/video-podcast-maker

    Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection →…

    1.7k GitHub stars~3.6k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Muapi Director

    Anil-matcha/vox-ai-motion-graphics-generator

    Turn ONE topic, talking-head video, or photo into a finished Vox-style paper-collage explainer / ad video on the MuAPI platform (api.muapi.ai) + local ffmpeg — script, collage keyframes, motion…

    246 GitHub stars~679 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • A skill your agent uses when building a local Windows WebCode installer from this repo for machine testing, especially when the package must bundle the Kokoro or sherpa-onnx Reply TTS service, model…

    278 GitHub stars~787 tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed

More from dtsola/xiaoyaosearch

  • Better Icons

    dtsola/xiaoyaosearch

    Searches more than 200 Iconify icon libraries and fetches icons as SVG from a command line tool or an MCP server.

    1k GitHub starsUsed in 2 repos~895 tokens
    Auto-check passed
  • Excalidraw Diagram Generator

    dtsola/xiaoyaosearch

    Writes .excalidraw JSON diagrams and exports them to PNG or SVG through the Kroki API or a local Firefox-based CLI.

    1k GitHub stars~3.5k tokensUpdated 9 days ago
    Auto-check passed
  • Searches local files from an agent through the XiaoyaoSearch MCP server, using semantic, full-text, hybrid, image-similarity and voice queries.

    1k GitHub stars~836 tokensUpdated 9 days ago
    Auto-check passed

Questions about Video Podcast Maker

What does Video Podcast Maker do?

Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references. Given a topic, the skill runs an automated pipeline from research and script through TTS narration and Remotion composition to an MP4 with background music, aimed at horizontal knowledge videos for Bilibili in either 4K or 1080p with a clean white style by default. The tech stack is the coding agent, a TTS backend, Remotion and FFmpeg, and it depends on a separate Remotion best-practices skill being loaded first.

When should I use Video Podcast Maker?

Video Podcast Maker fits situations like: making a narrated knowledge video from a single topic; learning visual design patterns from a reference video or image; producing a Bilibili-style horizontal video with background music; saving or reusing a video style profile.

How do I install Video Podcast Maker in Claude Code?

Run `npx skills add dtsola/xiaoyaosearch --skill video-podcast-maker -a claude-code`. Or copy the skill folder (.claude/skills/video-podcast-maker in dtsola/xiaoyaosearch) into .claude/skills/video-podcast-maker in your project. Claude Code loads it when a task matches its description.

How do I install Video Podcast Maker in Codex?

Run `npx skills add dtsola/xiaoyaosearch --skill video-podcast-maker -a codex`. Or copy the skill folder (.claude/skills/video-podcast-maker in dtsola/xiaoyaosearch) into .agents/skills/video-podcast-maker in your project. Codex loads it when a task matches its description.

Can I use Video Podcast Maker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dtsola/xiaoyaosearch --skill video-podcast-maker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-podcast-maker, .gemini/skills/video-podcast-maker, .github/skills/video-podcast-maker and .opencode/skills/video-podcast-maker in your project.

What does Video Podcast Maker need to run?

Going by SKILL.md and its folder, Video Podcast Maker needs the command-line tools its instructions call (git, npx and python3). Our summary lists: Remotion with the `remotion-best-practices` skill; FFmpeg; A TTS backend, with edge as the default; Python 3 for the prerequisites check.

Does Video Podcast Maker access the network?

SKILL.md contains no URLs. Its commands use git and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Video Podcast Maker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Podcast Maker use?

Video Podcast Maker is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Podcast Maker use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.

What are the alternatives to Video Podcast Maker?

Skills that share tags, products or a category with Video Podcast Maker: Vox Director (Alisa0808/vox-director, 2.2k stars), Whiteboard Video (gnipbao/codex-whiteboard-video-skill, 327 stars), Super Video Maker (Bomx/super-video-maker-skill, 310 stars) and Video Podcast Maker Nano (Agents365-ai/video-podcast-maker, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Podcast Maker?

dtsola (a GitHub user) maintains it in dtsola/xiaoyaosearch, which has 1,036 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on October 2, 2026.

Source: dtsola/xiaoyaosearch on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.