Agent skill

Video Podcast Maker Lite

by Agents365-ai in Agents365-ai/video-podcast-maker

Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion.

MITAuto-check passedMedia & Creative

Install Video Podcast Maker Lite

skills CLI
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-lite -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker-lite --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-podcast-maker-lite .claude/skills/video-podcast-maker-lite && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-podcast-maker-lite
GitHub stars
1.7k
Token cost
~4.5k tokens
SKILL.md length
2,109 words
Files
3 (incl. scripts, references)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion.

  • Works in 6 steps: Write the script:… → TTS → Compose visuals → …
  • The user wants a quick narrated video from a topic without the full video-podcast-maker machinery (no extra skills
  • SKILL.md covers Prerequisites, Project discovery (run before…, Workflow (6 steps) and Composition contract, plus 3 more sections
  • Runs Python scripts from its folder; calls npx, pip3 and ffmpeg; reaches pixabay.com; needs AZURE_SPEECH_KEY

What it does

Video Podcast Maker Lite is an agent skill from Agents365-ai/video-podcast-maker. Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion. Use when the user wants a quick narrated video from a topic without the full video-podcast-maker machinery (no extra skills, no thumbnails/shorts/publish matrix). Do NOT trigger for heavy production needs — use video-podcast-maker for those.

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/pronunciation.md` and `scripts/tts.py`).

It sits in Media & Creative, covering Text to speech and voice, Video production and AI video generation. It works with Remotion, Microsoft Azure and Bilibili. The repository describes itself as: Topic → 4K narrated video for coding agents. v5.3.0: local TTS (edge free + azure, no external engine), manifest-based Asset Engine, Remotion composition, cost-gated AI…. The licence is MIT.

When your agent uses it

  • The user wants a quick narrated video from a topic without the full video-podcast-maker machinery (no extra skills
  • No thumbnails/shorts/publish matrix)
  • Heavy production needs — use video-podcast-maker for those

Example prompts

  • “/video-podcast-maker-lite”

Requirements

  • Python 3
  • Node.js
  • A credential in AZURE_SPEECH_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Write the script: videos/{name}/podcast.txt
  2. TTS
  3. Compose visuals
  4. Preview (mandatory human gate)
  5. Render + BGM
  6. Delivery extras (only when the user's pipeline includes them)

What it can do on your machine

Read from SKILL.md and the folder at commit 33b8078. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • pip3
    • ffmpeg
    • npm
    • python3
    • curl
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • pixabay.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_SPEECH_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Podcast Maker Lite loads about 4.5k tokens when it runs, and up to ~5.3k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 2,109 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Agents365-ai/video-podcast-maker at commit 33b8078, republished under its MIT licence (© Agents365-ai). 2,109 words, ~4,498 tokens.

Download SKILL.mdSave it as .claude/skills/video-podcast-maker-lite/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
video-podcast-maker-lite
description
Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion. Use when the user wants a quick narrated video from a topic without the full video-podcast-maker machinery (no extra skills, no thumbnails/shorts/publish matrix). Do NOT trigger for heavy production needs — use video-podcast-maker for those.
argument-hint
[topic]
author
Agents365-ai
category
Content Creation
version
1.1.0

Video Podcast Maker Lite

A single-purpose pipeline for personal use: topic → narration script → Azure TTS (SSML) → Remotion → 1080p/4K MP4. No external skills, no config files, no bundled templates — the Remotion composition is generated once against the contract below, then reused across videos.

Prerequisites

bash
pip3 install azure-cognitiveservices-speech   # the only Python dependency
export AZURE_SPEECH_KEY="..."      # Azure Speech resource key
export AZURE_SPEECH_REGION="..."   # e.g. eastasia
# ffmpeg + node 18+ required; one Remotion project with npm install done
# Playwright MCP (session browser) needed only for Step 5 BGM fetching

Project discovery (run before Step 1)

  1. Locate the Remotion project — the directory containing both src/remotion/index.ts and node_modules/remotion/. Check the current working directory first; if it is not there, ask the user for the project path. Do NOT scan the filesystem, and do NOT create a new project if one exists elsewhere.
  2. Inventory existing videos — ls videos/ and ls src/remotion/*Video.tsx. The most recent *Video.tsx is the component to copy in Step 3; videos/ shows which videos already exist.
  3. Existing videos/{name}/? Then this is an iteration on that video, not a new one — reuse the directory and re-run only what changed (see Iterating).
  4. No project anywhere (first run only) → scaffold once in the working directory: npm init -y && npm i remotion @remotion/cli @remotion/transitions react react-dom, create src/remotion/, and warn the user about the one-time ~2 GB install. All later videos skip this.

Workflow (6 steps)

Run all commands from the project root. All artifacts for one video live in videos/{name}/ inside the project. {name} is lowercase English, hyphen-separated.

Step 1 — Write the script: videos/{name}/podcast.txt

One [SECTION:xxx] marker per video segment; section names are lowercase English (hero, content-1, outro). An optional display title goes after a | — [SECTION:outro|thanks] — it labels the progress-bar pill and default layout (without it, the label derives from the first sentence, which can be awkward). Lines starting with # are ignored (and may safely mention markers). Spoken text only — no markdown — and follow the script style rules below. Example:

text
# comment lines are not spoken
[SECTION:hero|intro]
欢迎来到本期视频!今天我们要聊一个大家都关心的话题。

[SECTION:content-1|point 1]
首先,我们来看第一个要点。这里有几个关键信息需要你知道。

[SECTION:outro|thanks]
好了,今天的内容就到这里。如果觉得有帮助,欢迎点赞关注,我们下期再见!
Script style (anti-AI-flavor, zh-CN)

Apply while writing, then self-check before TTS. Goal: everyday spoken Chinese, not written prose with commas. (Distilled from the full skill's natural-narration.md + script-polish.md — those are the canonical sources; edit rules there first, then mirror here.)

Connector swap (written → spoken): 此外→还有 · 然而→但是 · 因此→所以 · 与此同时→这时候 · 总的来说/综上所述→删掉 · 首先/其次/最后→直接讲下一件事。

Kill list (rewrite or delete): 赋能、打造、深入探讨、值得一提/值得注意的是、众所周知、至关重要、革命性、颠覆、天花板、无缝、闭环、抓手、里程碑、标志着、未来可期、让我们拭目以待。

Structural tells — the fix patterns:

PatternFix
Verb-noun shells: 进行优化/实现增长/做出选择concrete action + result: "把审批从三步改成一步"
Negation contrast: 不是 X,而是 Ystate Y directly
Three-part parallelism: 既是…又是…更是…keep the most informative item; two beat three
Empty intensifiers: 显著/大幅/非常a number or a perceivable consequence
Vague attribution: 业内普遍认为/有专家指出named source + date, else delete the sentence
Slogan endings: 未来可期/注入新的活力land on a concrete fact: number, date, next action

Write for the ear: one idea per sentence, subject first, vary sentence length, no nested clauses, no — or · as connectives (they don't get spoken and clutter subtitles). A light first person is fine ("我实测下来").

Subtitles are the script, verbatim — so write numbers the way they should LOOK on screen: Arabic digits (3.8, 63K, 98.8%, 128G), never Chinese numerals (六十五点一 / 三百九十七). The same digit rule applies to on-screen text in the Remotion components (cards, headlines). Do NOT write the spoken form into the script to fix pronunciation — it leaks into subtitles. tts.py derives the spoken layer itself: every number-bearing token (86.1, 9B, 5600, Qwen3.5) is converted to its Chinese reading (八十六点一, 九B, 五千六百, 千问三点五) before synthesis, then word boundaries are mapped back so subtitles keep the display text. (Multilingual voices like zh-CN-XiaoxiaoMultilingualNeural read bare digits in English in mixed context — that is exactly what this layer prevents. SSML <sub alias> was tried and abandoned: Azure's word-boundary events for <sub> are buggy and corrupt the SRT.)

Numbers must be traceable — a precise number without a source is fabricated; drop it or attribute it.

Self-check before Step 2: no kill-list words? no "不是…而是…"? no slogan ending? Read each section aloud — if you stumble, split the sentence.

STOP — script review gate (mandatory): when podcast.txt is written, halt the pipeline and hand the script to the user for review. Do NOT run TTS (Step 2) until the user explicitly approves the script. This gate comes before everything else downstream — audio, timings, and visual entrances all derive from the script, so a late script change costs a full re-run.

Step 2 — TTS
bash
python3 "${SKILL_DIR}/scripts/tts.py" videos/{name}/podcast.txt videos/{name}/

Run --check first (lint-only, no synthesis, free) to catch polyphone and alias gaps before paying for synthesis.

Produces podcast_audio.wav + podcast_audio.srt + timing.json + cues.json (per-cue text, global frame, and section_frame — use it to align visual entrances in the component instead of hand-parsing the srt). Each section is synthesized separately via SSML and concatenated, so section timings are exact by construction. Subtitle cues are phrase-first: a sentence within 30 visible chars is shown whole; longer sentences are packed from comma/semicolon clauses (~22 per cue); an over-long clause is cut at the word boundary nearest its middle, never mid-word; tiny trailing cues merge into the previous one within the same section only (a short first sentence of the next section must never bleed into the previous section's last cue). Re-running over an existing timing.json prints a per-section duration diff; any section that moved >0.3s means the component's hardcoded entrance delays need re-aligning from the new cues.json.

Knobs: --voice (default zh-CN-XiaoxiaoNeural), --style (mstts:express-as, e.g. gentle / cheerful — stick to these two; others can sound strained), --rate (prosody, e.g. -4%), --phonemes (whole-word pinyin dict for polyphones like 命令行/同行; by default ~/.video-podcast-maker/phonemes.json and phonemes.json next to the input are merged — per-video entries win), --aliases (pronunciation aliases display→spoken, e.g. "Ornith-1.5": "Ornith 一点五"; same merge order with aliases.json). Env fallbacks: TTS_VOICE / TTS_STYLE / TTS_RATE. For a consistent channel voice across videos, set them once in your shell profile (e.g. TTS_VOICE=zh-CN-XiaoxiaoMultilingualNeural TTS_RATE=+5%) instead of passing flags every time.

Pronunciation: polyphone pre-flight, known misreading shapes, alias-dict mechanics, and built-in spoken-layer behaviors are documented in references/pronunciation.md — read it before the first TTS run of a new topic.

Step 3 — Compose visuals
  • First video ever: generate the composition (index.ts + Root.tsx + Video.tsx under src/remotion/) against the composition contract.
  • Every later video: copy the previous video's component and edit — never start from the contract again.

One Remotion project hosts all videos; a new video adds exactly one component file and one <Composition> registration:

text
project-root/                    # ONE project, npm install once
├── src/remotion/
│   ├── index.ts                 # registerRoot — shared, never changes
│   ├── Root.tsx                 # one <Composition id="…"> per video
│   ├── DemoVideo.tsx            # per-video component (copy of the last one, edited)
│   └── NextTopicVideo.tsx
└── videos/{name}/               # per-video artifacts: podcast.txt, wav, srt, timing.json, mp4

Per video: pick a unique PascalCase component/composition id (e.g. ReferenceManagerComparison), set title / colors, and give each section name a layout (a switch (section.name) over hero / content-N / outro works well).

Visual richness (default style — plain text-in-a-box is not the target look): every info card, flow box, pill and stat tile carries ONE leading emoji (or an @lobehub/icons brand component when it depicts a real product); at most one per element, and never inside a MONO value line (emoji break monospace alignment — put it on the label/title instead). Keep the mapping one concept = one emoji consistent across the whole video and pick it tastefully for the topic (dates/parameters/speed/cost each get an obvious match; the agent chooses). Every section also gets at least one visual anchor (an illustration): official material first (product banner, spec card, screenshot), else a free illustration or icon set (unDraw SVG, Pixabay/Pexels images, OpenMoji / Microsoft Fluent Emoji / Google Noto Emoji, @lobehub/icons for brand logos) — note the source + license in the publish_info asset-sources section (attribution-required sets like Flaticon's free tier must be credited in the video description). Emoji decorate the UI cards, illustrations anchor the section; neither replaces the other.

Step 4 — Preview (mandatory human gate)
bash
npx remotion studio src/remotion/index.ts --public-dir videos/{name}/

MUST launch Studio and wait for the user to review in person. NEVER render until the user explicitly confirms ("渲染" / "render"). An adjustment request is not confirmation — apply the change, let Studio hot-reload, and ask again. Every round of visual changes needs its own fresh confirmation; confirmation never carries over to Step 5.

Show full SKILL.md (886 more words)Show less
Step 5 — Render + BGM

Render:

bash
npx remotion render src/remotion/index.ts MyVideo videos/{name}/output.mp4 --public-dir videos/{name}/

BGM (default; skip only if the user says no music): fetch one random track from Pixabay Music and mix it at low volume (narration stays dominant). Pixabay License: royalty-free, commercial use OK, no attribution required — still log title/author into the publish_info asset-sources section.

How to fetch (verified 2026-08-25): Pixabay's search/list pages are client-rendered and Cloudflare-gated for non-browser clients, but a single-track page opened in a real browser embeds the full-track download URL in its JSON-LD, and that cdn.pixabay.com URL then downloads fine with plain curl.

  1. With the session's browser (Playwright MCP), open a search page — https://pixabay.com/music/search/cinematic/ (or relaxing / ambient if a calmer bed is wanted).

  2. Collect result links matching /music/<slug>-<id>/ (exclude /music/search/ and locale-prefixed ones like /de/music/...), pick one at random, open it.

  3. Read the track's JSON-LD (script[type=application/ld+json"] → the AudioObject): contentUrl (full track MP3), name, creator.name, duration. Prefer ~1.5–8 min; if out of range, open another link. If the page's JSON-LD has no contentUrl (layout changed), stop retrying — ask the user to pick a track and provide the download URL.

  4. Download with curl (browser UA + Referer: https://pixabay.com/ — verified to work):

    bash
    curl -sL -A "<browser UA>" -H "Referer: https://pixabay.com/" "<contentUrl>" -o videos/{name}/bgm.mp3

Mix (bgm low; the loudnorm stage is required — without it the mix lands ≈ -32 dB mean, ~10 dB too quiet):

bash
ffmpeg -y -i videos/{name}/output.mp4 -i videos/{name}/bgm.mp3 -filter_complex "[1:a]volume=0.08[bg];[0:a][bg]amix=inputs=2:duration=first[a];[a]loudnorm=I=-16:TP=-1.5:LRA=11[out]" -map 0:v -map "[out]" -c:v copy -shortest videos/{name}/final_video.mp4

Stop the Studio server once the render is confirmed — a leftover Studio holds its port and keeps watching files.

Step 6 — Delivery extras (only when the user's pipeline includes them)

Rendered after Step 5, in this order:

  • Cover stills: 16x9 + 4x3 Thumbnail stills via npx remotion still — delete the old PNG first, stills skip existing files.
  • BGM loudness check: ffmpeg -i final_video.mp4 -af volumedetect -f null -, mean ≈ -19 to -22 dB, max ≈ -1.5 dB.
  • publish_info.md: title / description / tags / chapter timestamps (chapters derive from timing.json).
  • assets/manifest.json: asset provenance.
  • Repo-level bookkeeping: index/theme generators (e.g. a VIDEO_INDEX.md builder) typically key off publish_info.md titles, and classification scripts may refuse to run until the new video dir is added to their assignment map — run them after the publishing artifacts are in place.

Composition contract

The composition consumes three files from --public-dir videos/{name}/ via staticFile(): podcast_audio.wav, podcast_audio.srt, and timing.json:

json
{
  "total_duration": 32.1,
  "fps": 30,
  "total_frames": 964,
  "sections": [{ "name": "hero", "label": "…", "start_time": 0, "duration": 11.8, "start_frame": 0, "duration_frames": 355 }]
}

Non-negotiables when generating a composition (each one is a real failure mode if missed):

  1. Audio is the master clock — calculateMetadata returns durationInFrames = timing.total_frames, loaded at runtime. Never hardcode a duration. Resolution per project convention: 1920×1080 @ 30fps, or 4K (3840×2160) via a wrapper that scales a 1920×1080 design ×2 (keep the inner design at 1080p coordinates).
  2. Async assets gate rendering — load timing.json/SRT with fetch(staticFile(...)) wrapped in delayRender/continueRender, or the first frames render without data/subtitles. When copying an existing component that instead bundles timing.json via a direct import, that convention is equally valid (the JSON is inlined at bundle time) — follow the copied component.
  3. Compensate TransitionSeries overlap — rendered length is sum(sections) − (N−1)×transitionFrames. Scale every section proportionally so the total lands exactly on total_frames (absorb rounding on the last section); do not pad the first section.
  4. Subtitles — parse the SRT, show the cue active at the current frame, positioned bottom-center above the progress bar (bottom ≈ 70px). Body text ≥ 32px, titles ≥ 64px.
  5. Fail loud — cancelRender with the real error if timing.json fails to load; a silent hang costs a render-timeout to diagnose.
  6. Props type must be a type alias, not an interface — Remotion constrains props to Record<string, unknown>, which interfaces fail (no implicit index signature). type VideoProps = { ... } passes.
  7. Chapter progress bar — pinned to the very bottom (height ~55px at 1080p): one pill per section with flex proportional to duration_frames and the section label as text (~24px); the active pill is filled with primaryColor and gets a translucent white overlay whose width is the intra-chapter progress, past pills gray, future pills outlined; plus a ~3px global progress line along the bottom edge. Section layouts keep the bottom ~200px clear in total (bar + subtitle zone).

Iterating

  • Script changed → re-run Step 2 (timestamps all shift — never hand-edit timing.json), then re-render.
  • Audio re-synthesized (same script, different voice/rate) → re-run Step 2, then visuals may stay if timings didn't shift; re-enter the Step 4 gate before rendering.
  • Visuals only → edit the video's component, re-render (audio untouched). This re-enters the Step 4 gate: fresh in-person confirmation required before rendering.
  • Reuse the same videos/{name}/ directory; never start a new project per video.

Rules

  1. Studio before render. Never render without a fresh, explicit in-person confirmation in the current Studio session (see Step 4).
  2. After rendering, output.mp4 duration must match podcast_audio.wav within ±0.5s (ffprobe both). If not, the composition contract (items 1/3) is violated — fix the composition, not the timing file.
  3. Always --public-dir videos/{name}/ on every Remotion command — it is how the composition finds timing.json, the WAV, and the SRT.
  4. One Remotion project for all videos (see the layout in Step 3); only videos/{name}/ and the per-video component change.

Troubleshooting

  • Azure Speech SDK is not installed → pip3 install azure-cognitiveservices-speech.
  • Set AZURE_SPEECH_KEY and AZURE_SPEECH_REGION first → export both env vars (see Prerequisites).
  • TTS network/auth failure → the script retries once per section; persistent CancellationReason.Error usually means a bad key/region or an unsupported voice/style combo.
  • ffmpeg: command not found → brew install ffmpeg.
  • Video/audio duration mismatch > 0.5s → re-run Step 2; if it persists, the composition is violating contract item 1 or 3.
  • timing.json not found in Studio/render → missing --public-dir videos/{name}/.

© Agents365-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in skills/video-podcast-maker-lite of Agents365-ai/video-podcast-maker.

  • SKILL.md
  • references/pronunciation.md
  • scripts/tts.py

Open the folder on GitHubat commit 33b8078

Compare with similar skills

Video Podcast Maker Lite next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Podcast Maker Lite compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Podcast Maker Lite this skillAgents365-ai/video-podcast-maker1.7k—~4.5kAutomated safety check: PassMIT
Anything2explainerVincentwei1021/anything2explainer2.4k—~2.7kAutomated safety check: PassCustom licence
Super Video MakerBomx/super-video-maker-skill310—~11kAutomated safety check: NotesNone
Suggest Sfxhassancs91/claude-youtube-editor328—~3kAutomated safety check: PassMIT
AI Mediaericrisco/rsc-harness190—~3.3kAutomated safety check: PassMIT
Media ProductionWrongStack/WrongStack371—~1kAutomated safety check: PassMIT

Similar skills

  • Anything2explainer

    Vincentwei1021/anything2explainer

    给一个主题,产出一条黑底 MG 风格(幕底可选星点或点阵波)、有配音字幕章节进度条的科普讲解视频(中文或英文;Remotion 代码动画;时长由用户定,常用 3–5 分钟)。内含可编译模板、图元库、配音/分镜/渲染工具、风格与动效规范、多 agent 分工协议与 QC 判据,以及一条完整样片(《RAG 与知识库》)作为质量标尺。Turn any topic into a narrated…

    2.4k GitHub stars~2.7k tokensUpdated 23 days ago
    Media & CreativeAuto-check passed
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    310 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Suggest Sfx

    hassancs91/claude-youtube-editor

    Step 4 of the AI Video Editor pipeline — the SFX pass. An agent skill from hassancs91/claude-youtube-editor.

    328 GitHub stars~3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • AI Media

    ericrisco/rsc-harness

    A skill your agent uses when a creative goal must become a finished media file: pick and order generative-media models per modality — AI voiceover, image-to-video clips, score — then glue them with…

    190 GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Media Production

    WrongStack/WrongStack

    Create and process finished videos with Remotion, Motion Canvas, Manim, FFmpeg or an available AI video provider.

    371 GitHub stars~1k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Vocabulary Video Pipeline Ark Plan

    dracohu2025-cloud/draco-skills-collection

    【Ark Agent Plan 专用版本】基于 Remotion 的英文词汇视频自动化生成流水线。输入一个英文单词,自动完成:诊断、火山引擎 TTS 音频(与 Seedream/Seedance 共享认证)、节奏分割、视频渲染、飞书上传和成本汇报。

    227 GitHub stars~1k tokensUpdated 24 days ago
    Media & CreativeAuto-check: notes

More from Agents365-ai/video-podcast-maker

  • Video Podcast Maker

    Agents365-ai/video-podcast-maker

    A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat…

    1.7k GitHub stars~4.9k tokensUpdated 10 days ago
    Auto-check passed
  • Video Podcast Maker Nano

    Agents365-ai/video-podcast-maker

    Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection →…

    1.7k GitHub stars~3.6k tokensUpdated 10 days ago
    Auto-check passed

Questions about Video Podcast Maker Lite

What does Video Podcast Maker Lite do?

Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion. Video Podcast Maker Lite is an agent skill from Agents365-ai/video-podcast-maker. Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion.

When should I use Video Podcast Maker Lite?

Video Podcast Maker Lite fits situations like: the user wants a quick narrated video from a topic without the full video-podcast-maker machinery (no extra skills; no thumbnails/shorts/publish matrix); heavy production needs — use video-podcast-maker for those.

How do I install Video Podcast Maker Lite in Claude Code?

Run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-lite -a claude-code`. Or copy the skill folder (skills/video-podcast-maker-lite in Agents365-ai/video-podcast-maker) into .claude/skills/video-podcast-maker-lite in your project. Claude Code loads it when a task matches its description.

How do I install Video Podcast Maker Lite in Codex?

Run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-lite -a codex`. Or copy the skill folder (skills/video-podcast-maker-lite in Agents365-ai/video-podcast-maker) into .agents/skills/video-podcast-maker-lite in your project. Codex loads it when a task matches its description.

Can I use Video Podcast Maker Lite in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-lite -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-podcast-maker-lite, .gemini/skills/video-podcast-maker-lite, .github/skills/video-podcast-maker-lite and .opencode/skills/video-podcast-maker-lite in your project.

What does Video Podcast Maker Lite need to run?

Going by SKILL.md and its folder, Video Podcast Maker Lite needs Python for the scripts in its folder, the command-line tools its instructions call (npx, pip3, ffmpeg, npm, python3 and curl) and credentials named AZURE_SPEECH_KEY. Our summary lists: Python 3; Node.js; A credential in AZURE_SPEECH_KEY.

Does Video Podcast Maker Lite access the network?

SKILL.md names 1 domain. In commands or code: pixabay.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Video Podcast Maker Lite safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Podcast Maker Lite use?

Video Podcast Maker Lite is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Podcast Maker Lite use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 772 tokens, read only when the agent opens those files.

What are the alternatives to Video Podcast Maker Lite?

Skills that share tags, products or a category with Video Podcast Maker Lite: Anything2explainer (Vincentwei1021/anything2explainer, 2.4k stars), Super Video Maker (Bomx/super-video-maker-skill, 310 stars), Suggest Sfx (hassancs91/claude-youtube-editor, 328 stars) and AI Media (ericrisco/rsc-harness, 190 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Podcast Maker Lite?

Agents365-ai (a GitHub user) maintains it in Agents365-ai/video-podcast-maker, which has 1,670 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 1, 2026.

Source: Agents365-ai/video-podcast-maker on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.