Agent skill

Video Podcast Maker

by Agents365-ai in Agents365-ai/video-podcast-maker

A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat…

MITAuto-check passedMedia & Creative

Install Video Podcast Maker

skills CLI
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-podcast-maker .claude/skills/video-podcast-maker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-podcast-maker
GitHub stars
1.7k
Token cost
~4.9k tokens
SKILL.md length
1,610 words
Files
109 (incl. scripts, references, assets)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat…

  • Works in 4 steps: Audio is the master clock. Every slide… → Generate timing from TTS, not from text… → Never hand-write timing.json before… → …
  • The user gives a topic and wants an automated topic-driven narrated explainer
  • SKILL.md covers Contents, Bootstrap, Execution Modes and Regenerating an Existing Video, plus 7 more sections
  • Calls npx, python3 and git; needs AZURE_SPEECH_KEY

What it does

Video Podcast Maker is an agent skill from Agents365-ai/video-podcast-maker. Use when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels), or asks to learn visual design patterns from a reference video/image. Trigger when the user mentions creating a knowledge video, narrated explainer, video podcast, or animated infographic-style video from a topic — even if they don't say "video podcast" explicitly. Also trigger when the user wants to regenerate, re-render…

Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 110 other files, including scripts, reference files and assets (for example `.skillspector-baseline.yaml`, `AGENTS.md` and `package.json`).

It sits in Media & Creative, covering Video production and Text to speech and voice. It works with Remotion, Bilibili, Douyin and WeChat. The repository describes itself as: Topic → 4K narrated video for coding agents. v5.3.0: local TTS (edge free + azure, no external engine), manifest-based Asset Engine, Remotion composition, cost-gated AI…. The licence is MIT.

When your agent uses it

  • The user gives a topic and wants an automated topic-driven narrated explainer
  • Knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels)
  • Asks to learn visual design patterns from a reference video/image
  • The user mentions creating a knowledge video

Example prompts

  • “t say”
  • “/video-podcast-maker”

Requirements

  • Python 3
  • Node.js
  • A credential in AZURE_SPEECH_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Audio is the master clock. Every slide start, subtitle, chapter, and animation beat is derived from podcast_audio.wav and podcast_audio.srt.
  2. Generate timing from TTS, not from text estimates. Pipeline: podcast.txt → generate_tts.py → podcast_audio.wav + podcast_audio.srt +…
  3. Never hand-write timing.json before audio exists. If you already have curated slides, run align_timing_from_srt.py to anchor them to the…
  4. Compensate TransitionSeries overlap. TransitionSeries renders sum(section.duration_frames) - (N-1) * transitionFrames frames. Scale every…

What it can do on your machine

Read from SKILL.md and the folder at commit 33b8078. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • python3
    • git
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com
    • remotion.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_SPEECH_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Podcast Maker loads about 4.9k tokens when it runs, and up to ~43k if it reads all its reference files. Until then it costs about 248 tokens; SKILL.md has 1,610 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~248
When it runs · the whole SKILL.md, loaded when a task matches
~4.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~43k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Agents365-ai/video-podcast-maker at commit 33b8078, republished under its MIT licence (© Agents365-ai). 1,610 words, ~4,870 tokens.

Download SKILL.mdSave it as .claude/skills/video-podcast-maker/SKILL.md (or your agent's skills folder). This skill also uses 108 other files; get the full folder from GitHub.
name
video-podcast-maker
description
Use when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels), or asks to learn visual design patterns from a reference video/image. Trigger when the user mentions creating a knowledge video, narrated explainer, video podcast, or animated infographic-style video from a topic — even if they don't say "video podcast" explicitly. Also trigger when the user wants to regenerate, re-render, rebuild, update, or iterate on a narrated video this skill already produced — e.g. they edited the script/prompt, changed the visuals, or swapped the background music and want the final video remade (reuse the existing videos/{name}/ directory, never start a new project). Do NOT trigger for generic video editing, trimming, format conversion, color grading, or non-narrative video tasks. Produces 4K video via research → script → TTS → Remotion → MP4 + BGM.
argument-hint
[topic]
effort
high
author
Agents365-ai
category
Content Creation
version
5.3.0
created
2025-01-27
updated
2026-07-30
permissions
env, file_read, file_write, network, shell
bilibili
https://space.bilibili.com/441831884
github
https://github.com/Agents365-ai/video-podcast-maker

Recommended: Load Remotion Best Practices

This skill benefits from remotion-best-practices (not bundled) for the full Remotion pattern library. It is optional — minimum rules are below if absent.

  • Pi: read the loaded skill at remotion-best-practices (listed in available skills).
  • Claude Code: invoke remotion-best-practices skill/tool before proceeding.

Not installed? Get it from remotion-dev/skills (docs: remotion.dev/docs/ai/skills).

If remotion-best-practices is not installed, minimum rules: chromium must be available, always wrap 4K content in <Scale4K>, use <TransitionSeries> with linearTiming, and treat audio as the master clock.

Video Podcast Maker

Automated pipeline for 4K Bilibili horizontal knowledge videos from a topic. Coding agent + TTS backend + Remotion + FFmpeg.

Contents


Bootstrap

Resolve SKILL_DIR to the directory containing this SKILL.md:

  • Pi: the agent knows the skill path from the loaded skill list — set SKILL_DIR to that directory before running commands.
  • Claude Code: ${CLAUDE_SKILL_DIR} is auto-populated.
bash
SKILL_DIR="${SKILL_DIR:-${CLAUDE_SKILL_DIR}}"

# Prerequisites (CLIs + backend env vars)
python3 "${SKILL_DIR}/scripts/check_prereqs.py"

Updates flow through the skills CLI (npx skills update video-podcast-maker -g); direct git-clone installs use git pull per the README. This skill performs no update checks.

Prereqs failures — see README.md for setup. The check is backend-aware (resolves TTS_BACKEND env → user_prefs.json global.tts.backend → edge default), so only env vars required by the active backend are validated.

First video in a new project? Prefer reusing an existing Remotion project with node_modules/ already installed — creating a fresh project downloads ~2.2 GB of npm packages plus a 90 MB Chrome headless shell (one-time per project). If the user has a project from a previous video, use it. If a fresh project is necessary, run npm install in the background while you do Steps 1-4 (topic research and script writing).

All rendering goes into videos/{name}/ — every output.mp4, final_video.mp4, and thumbnail_*.png lands directly in the per-video directory. Never render to an out/ or dist/ directory; the --public-dir videos/{name}/ convention keeps everything self-contained.

TTS engine — two local backends, no external skill:

  • edge (default) — free, no key, via edge-tts.
  • azure — needs AZURE_SPEECH_KEY + AZURE_SPEECH_REGION (Microsoft Speech SDK).

Each synthesizes in-house (scripts/tts/backends/native.py) — pronunciation (display → spoken → back to display for subtitles) and phoneme application are built in. check_prereqs.py validates the active backend's env vars.

Multi-platform TTS? The former ttscn component skill that provided the 9-backend matrix is no longer a dependency of this skill. If you want those platforms, install Agents365-ai/ttsCN separately and call it directly — this skill ships only edge + azure.

Design Learning shortcut: If the user provides a reference video/image or asks to save/list/delete style profiles, see references/design-learning.md instead of running the workflow below.


Execution Modes

Detect Auto Mode (default) vs Interactive Mode at workflow start — the Auto-default decision table and per-request overrides are in references/workflow-script.md.


Regenerating an Existing Video

If videos/{name}/ already exists and the user is iterating on a finished or in-progress video, reuse that directory. Do NOT start a new project or a new videos/{newname}/.

Pick the smallest re-run for what actually changed:

ChangedRe-runReuses (don't redo)
Narration script (podcast.txt)Step 7 (TTS) → Step 8 preview → render+mixtopic research + section design
Visuals only (components, layout, colors)Step 8 preview → render+mixaudio (podcast_audio.wav / timing.json)
Background music onlyRe-mix BGMoutput.mp4 (no re-render)
Subtitles onlyStep 10.1 finalizeoutput.mp4 / video_with_bgm.mp4

Any re-run that changes what the viewer sees or hears re-enters the Step 8 gate: apply the change, let Studio hot-reload, and wait for a fresh explicit "render 4K" — the previous confirmation does not carry over. A script change shifts every downstream timestamp, so always regenerate timing.json through TTS — never hand-edit it. After any re-run, re-verify:

bash
python3 ${SKILL_DIR}/scripts/verify_output.py videos/{name}/

Workflow

Iterating on a finished video? If videos/{name}/ already exists, see Regenerating an Existing Video above for the minimal re-run — do NOT start at Step 1.

At Step 1 start, create one task per step in your agent's tracker. Mark in_progress on start, completed on finish. Files in videos/{name}/ are the durable record — if interrupted, inspect the directory to determine where to resume.

#StepOutputPhase file
1Define topic directiontopic_definition.mdworkflow-script.md
2Research topictopic_research.mdworkflow-script.md
3Design 5-7 sections(in-memory)workflow-script.md
4Write narration scriptpodcast.txtworkflow-script.md
4.5Pronunciation pre-flight (zh-CN)phonemes.jsonworkflow-script.md
5Asset plan & resolveassets/manifest.jsonworkflow-assets.md
6Generate thumbnails (16:9 + 4:3)thumbnail_*.pngworkflow-production.md
7Generate TTS audiopodcast_audio.wav, timing.jsonworkflow-production.md
8Remotion composition + Studio preview—workflow-production.md
9Render 4K + mix BGMoutput.mp4, video_with_bgm.mp4workflow-production.md
10Publish info + verify outputpublish_info.md, final_video.mp4workflow-publish.md
11Generate vertical shorts (optional)shorts/workflow-publish.md

Mandatory stops (bold rows above):

  • Step 8 — Studio review. MUST launch npx remotion studio and wait for user feedback before rendering. NEVER render 4K until the user explicitly confirms ("render 4K" / "render final"). A reply containing adjustment requests is not confirmation — apply the changes, let Studio hot-reload, and ask again. Every round of adjustments needs its own fresh confirmation before Step 9.
  • Step 10 — verify_output.py. MUST pass before declaring the video done. Exit 0 = green; exit 2 = warnings still publishable. Auto-fixes common omissions (creates final_video.mp4 if missing). Validates publish info (title, description, tags, chapters) against the platform matrix — generate it in Steps 5.5 and 10.2. For machine-readable output add --format json.

Pre-render audit (recommended) — before Step 8:

bash
python3 ${SKILL_DIR}/scripts/audit_beat_sync.py <Video.tsx> <timing.json>

Flags beats that drift > 1.5s from narration.

Auto Mode: visual self-review. When running in Auto Mode (no user watching Studio), render 3-5 key frame stills before asking for render confirmation:

bash
npx remotion still src/remotion/index.ts <CompositionId> videos/{name}/_review_001.png --public-dir videos/{name}/ --frame=<midpoint_frame>

Pick frames at: hero title (~10% in), a dense section midpoint, and the outro. Read the stills back as images and run the design-guide.md and visual-taste.md checklists against actual rendered output. Catch overflow, contrast, and layout regressions before the 4K render. Delete _review_*.png after review.

Show full SKILL.md (659 more words)Show less
Validation Checkpoints
After StepCheck
7 (TTS)podcast_audio.wav plays · timing.json covers all sections · SRT is UTF-8
9 (Render)output.mp4 is 3840×2160 · audio-video sync · no black frames
10 (Verify)verify_output.py exits 0 (or 2 with reviewed warnings)

Hard Rules

RuleRequirement
Single ProjectAll videos under videos/{name}/ in user's Remotion project. NEVER create a new project per video.
4K Output3840×2160 (or 2160×3840 vertical), use scale(2) wrapper over 1920×1080 design space
Audio SyncAudio (podcast_audio.wav + podcast_audio.srt) is the master clock. timing.json MUST be generated from the real TTS output, never hand-estimated. Before rendering, final video duration must match audio within ±0.5s. See Audio-Master Clock & Sync.
ThumbnailMUST generate both 16:9 (1920×1080) AND 4:3 (1200×900) — see design-guide.md
Studio Before RenderMUST launch remotion studio for review. NEVER render 4K until user explicitly confirms. Adjustment feedback ≠ confirmation — apply, hot-reload, ask again.
--public-dirEvery Remotion command uses --public-dir videos/{name}/. All output files (output.mp4, final_video.mp4, thumbnails) go directly into videos/{name}/ — never an out/ or dist/ dir.

Visual minimums (text sizes, content width, safe zones, animation safety) live in references/design-guide.md. MUST load before Step 8.

Audio-Master Clock & Sync

Golden rules
  1. Audio is the master clock. Every slide start, subtitle, chapter, and animation beat is derived from podcast_audio.wav and podcast_audio.srt.
  2. Generate timing from TTS, not from text estimates. Pipeline: podcast.txt → generate_tts.py → podcast_audio.wav + podcast_audio.srt + timing.json → composition → render.
  3. Never hand-write timing.json before audio exists. If you already have curated slides, run align_timing_from_srt.py to anchor them to the real SRT.
  4. Compensate TransitionSeries overlap. TransitionSeries renders sum(section.duration_frames) - (N-1) * transitionFrames frames. Scale every section proportionally to keep the rendered length equal to timing.total_frames. Do not stuff all overlap frames into the first section. The corrected pattern is in templates/Video.tsx.
Mandatory sync checkpoints
WhenCheck
After Step 7 (TTS)timing.json.total_duration matches podcast_audio.wav within ±0.5s
Before renderVideo.tsx scales all sections for transition overlap
After renderfinal_video.mp4 duration matches podcast_audio.wav within ±0.5s
Step 10 (verify)verify_output.py exits 0 and reports green on audio/timing

If any checkpoint fails, stop. Do not publish.

Output Specs
ParameterHorizontal (16:9)Vertical (9:16)
Resolution3840×2160 (4K)2160×3840 (4K)
Frame rate30 fps30 fps
EncodingH.264, 16MbpsH.264, 16Mbps
AudioAAC, 192kbpsAAC, 192kbps
Duration1-15 min60-90s (highlight)

Per-Video Layout

project-root/                           # Remotion project root
├── src/remotion/                       # Remotion source (Root.tsx, compositions, index.ts)
├── videos/{video-name}/                # Per-video directory
│   ├── topic_definition.md             # Step 1
│   ├── topic_research.md               # Step 2
│   ├── podcast.txt                     # Step 4: narration script
│   ├── phonemes.json                   # Step 4.5: zh-CN pronunciation overrides
│   ├── assets/manifest.json            # Step 5: per-section asset registry
│   ├── publish_info.md                 # Step 10: title/description/tags
│   ├── podcast_audio.wav               # Step 7: TTS audio
│   ├── podcast_audio.srt               # Step 7: subtitles
│   ├── timing.json                     # Step 7: timeline (drives animations)
│   ├── thumbnail_*.png                 # Step 6
│   ├── output.mp4                      # Step 9: 4K render
│   ├── video_with_bgm.mp4              # Step 9: with BGM
│   ├── final_video.mp4                 # Step 10: final output
│   └── bgm.mp3                         # Background music
└── remotion.config.ts
--public-dir per video

Every Remotion command uses --public-dir videos/{name}/ — each video's assets stay in its own directory, enabling parallel renders:

bash
npx remotion studio src/remotion/index.ts --public-dir videos/{name}/
npx remotion render ... videos/{name}/output.mp4 --public-dir videos/{name}/ --video-bitrate 16M
npx remotion still ... videos/{name}/thumbnail.png --public-dir videos/{name}/
Naming
  • Video name {video-name}: lowercase English, hyphen-separated (e.g. reference-manager-comparison)
  • Section name {section}: lowercase English, underscore-separated, matches [SECTION:xxx]
  • Thumbnails (16:9 AND 4:3 both required): thumbnail_remotion_16x9.png + thumbnail_remotion_4x3.png (or _ai_ prefix for AI-generated)

Additional Resources

Load on demand — do NOT load all at once:

FileLoad when
references/workflow-script.mdSteps 1-4 (topic → script) + Execution Modes (Auto vs Interactive)
references/natural-narration.mdLoad before Step 4 script writing — anti-slop rules for spoken narration (kill list, structural tells, checklist)
references/script-polish.mdLoad after Step 4 draft is written — deep editing toolkit with 24 EN+ZH before/after patterns, evidence boundaries, quality rubrics
references/workflow-assets.mdStep 5, or when the user supplies images/clips or wants stock/AI media
references/workflow-assets.mdA section needs a data-chart/infographic animation beyond the component library (transparent overlay via Hyperframes)
references/workflow-production.mdSteps 5.5-9.5 (publish info draft → thumbnails → TTS → Remotion → render → BGM mix)
references/workflow-publish.mdSteps 10-11 (publish info, verify, shorts)
references/platform-matrix.mdPlatform-specific behavior (thumbnails, chapters, outro, publish info, shorts)
references/design-guide.mdMUST load before Step 8 — visual minimums, typography, animation safety
references/visual-taste.mdLoad before Step 8 alongside design-guide — design dials, anti-default rules, visual modes, section rhythm
references/design-learning.mdUser provides a reference video/image, or manages style profiles
references/troubleshooting.mdChoosing Azure voice/style, debugging hoarse/glitchy audio
references/troubleshooting.mdOn error, script/CLI discovery, or user asks about preferences/BGM
templates/presets/kinetic-typography/Bold type-driven preset (opinion / argument / declaration videos)

All scripts are reachable through one dispatcher — start with python3 ${SKILL_DIR}/scripts/cli.py --help; full routes and envelope error codes: references/troubleshooting.md.


User Preferences

Mutable state (user_prefs.json, phonemes.json) lives in ~/.video-podcast-maker/ — safe from skill updates. Auto-migrated from the skill directory on first run. Run "show preferences" to view, or "set X Y" to change. Full commands: references/troubleshooting.md.


Troubleshooting

See references/troubleshooting.md on errors, BGM options, preference learning, design-learning issues.

© Agents365-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 108 other files (scripts, references, assets) in skills/video-podcast-maker of Agents365-ai/video-podcast-maker.

  • SKILL.md
  • .env.example
  • .python-version
  • .skillspector-baseline.yaml
  • AGENTS.md
  • LICENSE
  • assets/bilibili-triple-black.mp4
  • assets/bilibili-triple-white.mp4
  • assets/perfect-beauty-191271.mp3
  • assets/snow-stevekaldes-piano-397491.mp3
  • package.json
  • phonemes.template.json
  • prefs_schema.json
  • references/design-guide.md
  • references/design-learning.md
  • references/natural-narration.md
  • references/platform-matrix.md
  • references/script-polish.md
  • references/troubleshooting.md
  • … and 90 more

Open the folder on GitHubat commit 33b8078

Compare with similar skills

Video Podcast Maker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Podcast Maker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Podcast Maker this skillAgents365-ai/video-podcast-maker1.7k—~4.9kAutomated safety check: PassMIT
Video Downloaderkangarooking/kangarooking-skills662—~8.3kAutomated safety check: PassNone
Ra Video Wash PipelinePluviobyte/rnskill1.6k—~2.4kAutomated safety check: PassCustom licence
Media To TranscriptbozhouDev/video-skills-toolkit150—~1.8kAutomated safety check: NotesMIT
Learning Notes Automationchubbyguan/chubbyskills1.2k—~898Automated safety check: PassMIT
Suggest Sfxhassancs91/claude-youtube-editor328—~3kAutomated safety check: PassMIT

Similar skills

  • Video Downloader

    kangarooking/kangarooking-skills

    Download or open videos and recover platform captions, audio transcripts, keyframes, screen text, visual facts, and editing observations as a plain multimodaltranscript.md.

    662 GitHub stars~8.3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Ra Video Wash Pipeline

    Pluviobyte/rnskill

    End-to-end Chinese video washing pipeline. An agent skill from Pluviobyte/rnskill.

    1.6k GitHub stars~2.4k tokensUpdated 20 days ago
    Media & CreativeAuto-check passed
  • Media To Transcript

    bozhouDev/video-skills-toolkit

    Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0.

    150 GitHub stars~1.8k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Learning Notes Automation

    chubbyguan/chubbyskills

    学习笔记自动化:视频/播客转录 → 知识点提取 → 闪卡生成 → 知识图谱更新。触发词:学习笔记、闪卡、Anki、知识提取、视频学习

    1.2k GitHub stars~898 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Suggest Sfx

    hassancs91/claude-youtube-editor

    Step 4 of the AI Video Editor pipeline — the SFX pass. An agent skill from hassancs91/claude-youtube-editor.

    328 GitHub stars~3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video Downloader

    idiotLeoLYJ/Daliu-Awesome-Skills

    Download videos from Douyin, Kuaishou, Xiaohongshu, and Bilibili by sharing link.

    140 GitHub stars~475 tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings

More from Agents365-ai/video-podcast-maker

  • Video Podcast Maker Lite

    Agents365-ai/video-podcast-maker

    Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion.

    1.7k GitHub stars~4.5k tokensUpdated 9 days ago
    Auto-check passed
  • Video Podcast Maker Nano

    Agents365-ai/video-podcast-maker

    Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection →…

    1.7k GitHub stars~3.6k tokensUpdated 9 days ago
    Auto-check passed

Questions about Video Podcast Maker

What does Video Podcast Maker do?

A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat…. Video Podcast Maker is an agent skill from Agents365-ai/video-podcast-maker. Use when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels), or asks to learn visual design patterns from a reference video/image.

When should I use Video Podcast Maker?

Video Podcast Maker fits situations like: the user gives a topic and wants an automated topic-driven narrated explainer; knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat Channels); asks to learn visual design patterns from a reference video/image; the user mentions creating a knowledge video.

How do I install Video Podcast Maker in Claude Code?

Run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker -a claude-code`. Or copy the skill folder (skills/video-podcast-maker in Agents365-ai/video-podcast-maker) into .claude/skills/video-podcast-maker in your project. Claude Code loads it when a task matches its description.

How do I install Video Podcast Maker in Codex?

Run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker -a codex`. Or copy the skill folder (skills/video-podcast-maker in Agents365-ai/video-podcast-maker) into .agents/skills/video-podcast-maker in your project. Codex loads it when a task matches its description.

Can I use Video Podcast Maker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-podcast-maker, .gemini/skills/video-podcast-maker, .github/skills/video-podcast-maker and .opencode/skills/video-podcast-maker in your project.

What does Video Podcast Maker need to run?

Going by SKILL.md and its folder, Video Podcast Maker needs the command-line tools its instructions call (npx, python3, git and npm) and credentials named AZURE_SPEECH_KEY. Our summary lists: Python 3; Node.js; A credential in AZURE_SPEECH_KEY.

Does Video Podcast Maker access the network?

SKILL.md names 2 domains. As links in the text: github.com and remotion.dev. This is read from the text; nothing was executed.

Is Video Podcast Maker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Podcast Maker use?

Video Podcast Maker is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Podcast Maker use?

About 4.9k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 39k tokens, read only when the agent opens those files.

What are the alternatives to Video Podcast Maker?

Skills that share tags, products or a category with Video Podcast Maker: Video Downloader (kangarooking/kangarooking-skills, 662 stars), Ra Video Wash Pipeline (Pluviobyte/rnskill, 1.6k stars), Media To Transcript (bozhouDev/video-skills-toolkit, 150 stars) and Learning Notes Automation (chubbyguan/chubbyskills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Podcast Maker?

Agents365-ai (a GitHub user) maintains it in Agents365-ai/video-podcast-maker, which has 1,670 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 1, 2026.

Source: Agents365-ai/video-podcast-maker on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.