Super Video Maker
Bomx/super-video-maker-skill
End-to-end AI video production skill for agentic frameworks.
Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection →…
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker-nano --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-podcast-maker-nano .claude/skills/video-podcast-maker-nano && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-podcast-maker-nano" agent skill from https://github.com/Agents365-ai/video-podcast-maker/tree/main/skills/video-podcast-maker-nano into .claude/skills/video-podcast-maker-nano/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-podcast-maker-nano", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Agents365-ai/video-podcast-maker/tree/main/skills/video-podcast-maker-nanoType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker-nano --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/video-podcast-maker-nano .agents/skills/video-podcast-maker-nano && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-podcast-maker-nano" agent skill from https://github.com/Agents365-ai/video-podcast-maker/tree/main/skills/video-podcast-maker-nano into .agents/skills/video-podcast-maker-nano/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-podcast-maker-nano", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker-nano --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/video-podcast-maker-nano .cursor/skills/video-podcast-maker-nano && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-podcast-maker-nano" agent skill from https://github.com/Agents365-ai/video-podcast-maker/tree/main/skills/video-podcast-maker-nano into .cursor/skills/video-podcast-maker-nano/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-podcast-maker-nano", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Agents365-ai/video-podcast-maker.git --path skills/video-podcast-maker-nano--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker-nano --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/video-podcast-maker-nano .gemini/skills/video-podcast-maker-nano && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-podcast-maker-nano" agent skill from https://github.com/Agents365-ai/video-podcast-maker/tree/main/skills/video-podcast-maker-nano into .gemini/skills/video-podcast-maker-nano/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-podcast-maker-nano", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker-nanoInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/video-podcast-maker-nano .github/skills/video-podcast-maker-nano && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-podcast-maker-nano" agent skill from https://github.com/Agents365-ai/video-podcast-maker/tree/main/skills/video-podcast-maker-nano into .github/skills/video-podcast-maker-nano/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-podcast-maker-nano", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Agents365-ai/video-podcast-maker video-podcast-maker-nano --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Agents365-ai/video-podcast-maker.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/video-podcast-maker-nano .opencode/skills/video-podcast-maker-nano && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-podcast-maker-nano" agent skill from https://github.com/Agents365-ai/video-podcast-maker/tree/main/skills/video-podcast-maker-nano into .opencode/skills/video-podcast-maker-nano/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-podcast-maker-nano", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-podcast-maker-nanoSmallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection →…
Video Podcast Maker Nano is an agent skill from Agents365-ai/video-podcast-maker. Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection → script → TTS → video → 4K render ∥ publish info + cover. The skill defines the pipeline logic and self-verified checkpoints; any TTS backend and any video tool (Remotion, HyperFrames, CapCut, ...) work, and how much human oversight to apply is set by the working project's AGENTS.md/CLAUDE.md, not here. Use when the user…
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `AGENTS.template.md`).
It sits in Media & Creative, covering Text to speech and voice, Video production and Podcasting. It works with Remotion, HeyGen and Bilibili. The repository describes itself as: Topic → 4K narrated video for coding agents. v5.3.0: local TTS (edge free + azure, no external engine), manifest-based Asset Engine, Remotion composition, cost-gated AI…. The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 33b8078. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
ffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Video Podcast Maker Nano loads about 3.6k tokens when it runs. Until then it costs about 185 tokens; SKILL.md has 1,858 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Agents365-ai/video-podcast-maker at commit 33b8078, republished under its MIT licence (© Agents365-ai). 1,858 words, ~3,600 tokens.
.claude/skills/video-podcast-maker-nano/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.A 7-step pipeline for personal use: research ∥ materials → script → TTS → audio checkpoint → video → preview checkpoint → 4K render ∥ publish kit. No bundled scripts, no hardcoded backends, no templates. The skill owns the logic and runs autonomously by default; the TTS backend and video tool are chosen per video (see Tool selection).
What makes this work regardless of tool choice — the three invariants:
ffprobe both). Visuals are cut to the audio, never the reverse.Default mode is autonomous: the pipeline runs end to end, with the agent performing every checkpoint's self-verification. A working project's AGENTS.md/CLAUDE.md can override per checkpoint with one line each:
Checkpoint 1 (script): human — halt after Step 2 until the user approves the script. Worth it: a late script change costs a full re-run.Checkpoint 2 (audio): human — the user listens to the full audio before visuals. Worth it when TTS misreadings are costly to catch later.Checkpoint 3 (preview): human — the user reviews the draft before render. Worth it for style-sensitive channels.Unmentioned checkpoints stay agent-verified. Project policy files start from AGENTS.template.md (next to this SKILL.md): copy it into the video project root as AGENTS.md (plus a CLAUDE.md copy for Claude Code) and fill in the tool bindings. The skill never uploads or publishes anything anywhere — the publish kit is files on disk; publishing is always a human act outside this pipeline.
Decide the TTS backend and the video tool once, before Step 3, in this priority order:
AGENTS.md/CLAUDE.md in the working project (see Oversight policy) IS the user's standing specification — use it, no scanning, no second-guessing. If the user names a tool in-session, that overrides the file. If the policy file is missing or still contains template placeholders (/absolute/path/to/...), resolve the bindings via rules 2–3, then write the resolved values back into the project policy file so the next run starts bound.research.md; surface it in the final summary. Never block waiting for a tool confirmation.Capability floor (applies to rules 1–3): the TTS choice must produce narration audio plus subtitle timing; the video choice must export a draft video file (for Checkpoint 3 verification) and 4K. Live preview is a bonus, never a substitute for the draft export. Auxiliary jobs need no selection: research uses the built-in web search, stills/cover the video tool's still export or any image tool, duration checks ffprobe. If ffmpeg/ffprobe are absent, report it and stop — the ±0.5s invariant is non-negotiable and is not skipped to keep a run alive.
All artifacts for one video live in videos/{name}/ ({name} = lowercase English, hyphen-separated). File names for audio/timing adapt to the chosen backend; the set is what matters:
videos/{name}/
├── research.md # Step 1 — facts + sources
├── podcast.txt # Step 2 — narration script
├── podcast_audio.wav # Step 3 — narration audio (name per backend)
├── podcast_audio.srt # Step 3 — subtitle timing (or the backend's equivalent)
├── assets/ # Step 1 — images/BGM + sources.md (source + license per asset)
├── video-project/ # Step 5 — whatever the video tool produces
├── final_4k.mp4 # Step 7 — 3840×2160 render
├── cover.png # Step 7 — video cover
└── publish_info.md # Step 7 — title / description / tags / chaptersEntry point check (before Step 1). Look for videos/{name}/ at the project root (or the directory the user names). If it already contains artifacts from an earlier session, this is an iteration — resume from the earliest step affected (see Iterating), do not re-run Step 1. Only start at Step 1 when no artifacts exist.
Two parallel outputs from one investigation pass (facts and assets come from the same sources):
Research → research.md. Investigate the topic (web search, papers, the user's pointers). Distill into research.md: facts, numbers, and their sources. Every number the script will claim must trace back here — a precise number without a source is fabricated; drop it or attribute it.
Materials → assets/. Collect everything the visuals and cover will consume into videos/{name}/assets/: per-section images/illustrations/screenshots, brand logos, BGM. Two sources, in priority order:
assets/, never reference them in place.@lobehub/icons for brand logos).Record each asset's source URL + license in assets/sources.md at collection time (attribution-required sets must be credited in the video description at Step 7). No suitable asset exists for a section? Do not fabricate a screenshot — fall back to a text-only layout for that section (or a generic free-license illustration), record the decision in assets/sources.md, and move on; asking the user is optional in human mode. Missing assets can be added any time before Step 5.
podcast.txt, then Checkpoint 1Spoken text only, no markdown. Split the script into segments with [SECTION:xxx|display-label] markers (lowercase English names, e.g. [SECTION:hero|intro]) — one section per video segment. This marker convention is the portable contract between script, TTS chunking, and visual layout; any TTS/video tool can consume it. Markers are structural metadata: never spoken, never shown as subtitle text — TTS and subtitles consume the text between markers only.
Style rules (language-agnostic): see Script style.
Checkpoint 1 — script self-review (mandatory). Verify the script against every Script style rule, then reconcile its sections against assets/: list sections with no matching asset, collect or fall back per the no-asset rule above, and note what goes without. In human mode (project policy), halt instead and hand the script over — do NOT run TTS until the user explicitly approves. Audio, timings, and visual entrances all derive from the script; a late script change costs a full re-run.
Run the chosen TTS backend (Azure, Edge, fish, minimax, ... — whatever the session picked). Produce:
Pronunciation hygiene before synthesizing, in any narration language: brand/term readings that can't be derived mechanically (Qwen read as its Chinese brand name, MoE spelled letter-by-letter) go into an alias/phoneme list per the backend's mechanism — never into the script text (it would leak into subtitles).
Agent self-verification (autonomous mode): ffprobe the audio duration against a rough estimate from the script (per-language speaking rate; use it only to catch gross errors like a silent or truncated file); verify every brand/term token in the script has an alias/phoneme entry (a coverage check — actual pronunciation is exactly what human-mode Checkpoint 2 is for); check for silent or clipped segments (silencedetect/volumedetect) that suggest synthesis failures. Fix alias-list gaps and re-synthesize if found. In human mode, the user listens to the full audio instead; misreadings → fix the alias list (not the script), re-synthesize, re-check.
Cut the visuals with the chosen tool (Remotion, HyperFrames, CapCut, ...) to the narration audio, using assets/ as the material pool. Section markers from podcast.txt drive the layout; subtitle cues drive text entrances. Keep the draft files in videos/{name}/.
Agent self-verification (autonomous mode): inspect the draft export (a seekable video file — the capability floor guarantees one): extract one frame per section (ffmpeg -ss <mid-section-time> -i draft.mp4 -frames:v 1) and view each for layout overflow, missing subtitles, broken asset references; verify draft duration matches the narration audio within ±0.5s. Fix and re-verify after every change. In human mode, the user reviews the draft in person (live preview if the tool has one, else the draft export); render only on explicit confirmation ("render"), and every round of changes needs fresh confirmation.
Run in parallel (the render is the long blocking job; the publish kit doesn't depend on it):
final_4k.mp4. Verify duration vs narration audio within ±0.5s before calling it done.publish_info.md (title / description / tags / chapter timestamps — each chapter starts at the SRT time of its section's first cue) and cover.png (generate from the video tool's still frame if available, else any image tool; may reuse assets/ material). The asset-sources section of publish_info.md copies from assets/sources.md.Provenance: the full skill's
video-podcast-maker/references/natural-narration.md(anti-AI-flavor) +script-polish.md(deep editing) are the canonical sources; this is the language-agnostic distillation. Edit rules there first, then mirror here — do not fork a rule and drift it.
The narration language is whatever the user's script is — this pipeline defaults to Chinese but the rules below apply in any language's spoken register. These rules are enforced at Checkpoint 1; read each section aloud — if you stumble, split the sentence.
86.1, 1.5G, 9B) — subtitles are the script verbatim, so write what should LOOK on screen. Never write the TTS spoken form into the script to fix a misreading; use the backend's alias/phoneme layer.research.md — a precise number without a source is fabricated; drop it or attribute it.videos/{name}/ directory.© Agents365-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/video-podcast-maker-nano of Agents365-ai/video-podcast-maker.
Open the folder on GitHubat commit 33b8078
Video Podcast Maker Nano next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Video Podcast Maker Nano this skillAgents365-ai/video-podcast-maker | 1.7k | — | ~3.6k | Automated safety check: Pass | MIT | |
| Super Video MakerBomx/super-video-maker-skill | 310 | — | ~11k | Automated safety check: Notes | None | |
| Ra Video Production DirectorPluviobyte/rnskill | 1.6k | — | ~7.4k | Automated safety check: Notes | Custom licence | |
| Video Podcast Makerdtsola/xiaoyaosearch | 1k | — | ~3.4k | Automated safety check: Pass | MIT | |
| Making Demo Videosnukeop/nuclear | 19k | — | ~892 | Automated safety check: Pass | AGPL-3.0 | |
| HyperFrames Video Entry Pointheygen-com/hyperframes | 60k | 3 repos | ~5.2k | Automated safety check: Pass | Apache-2.0 |
Bomx/super-video-maker-skill
End-to-end AI video production skill for agentic frameworks.
Pluviobyte/rnskill
End-to-end video production orchestration for the content-creation workspace.
dtsola/xiaoyaosearch
Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.
nukeop/nuclear
A skill your agent uses when making a demo, tutorial, or feature video of Nuclear.
heygen-com/hyperframes
Entry point for making, editing and rendering videos from HTML compositions with HyperFrames, routing each request to the right workflow.
calesthio/OpenMontage
Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.
Agents365-ai/video-podcast-maker
A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat…
Agents365-ai/video-podcast-maker
Minimal personal narrated-video pipeline — a topic becomes a talking-head-free explainer MP4 (1080p or 4K) via script → Azure TTS (SSML) → Remotion.
Categories
Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection →…. Video Podcast Maker Nano is an agent skill from Agents365-ai/video-podcast-maker. Smallest personal narrated-explainer-video pipeline (spoken narration over visuals, not an audio podcast), fully tool-agnostic and autonomous by default — topic → research ∥ asset collection → script → TTS → video → 4K render ∥ publish info + cover.
Video Podcast Maker Nano fits situations like: the user wants a quick personal narrated video with minimal steps; not they name the tool stack; audio-only podcasts; written episodic content.
Run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a claude-code`. Or copy the skill folder (skills/video-podcast-maker-nano in Agents365-ai/video-podcast-maker) into .claude/skills/video-podcast-maker-nano in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a codex`. Or copy the skill folder (skills/video-podcast-maker-nano in Agents365-ai/video-podcast-maker) into .agents/skills/video-podcast-maker-nano in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Agents365-ai/video-podcast-maker --skill video-podcast-maker-nano -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-podcast-maker-nano, .gemini/skills/video-podcast-maker-nano, .github/skills/video-podcast-maker-nano and .opencode/skills/video-podcast-maker-nano in your project.
Going by SKILL.md and its folder, Video Podcast Maker Nano needs the command-line tools its instructions call (ffmpeg).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Video Podcast Maker Nano is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Video Podcast Maker Nano: Super Video Maker (Bomx/super-video-maker-skill, 310 stars), Ra Video Production Director (Pluviobyte/rnskill, 1.6k stars), Video Podcast Maker (dtsola/xiaoyaosearch, 1k stars) and Making Demo Videos (nukeop/nuclear, 19k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Agents365-ai (a GitHub user) maintains it in Agents365-ai/video-podcast-maker, which has 1,670 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 1, 2026.
Source: Agents365-ai/video-podcast-maker on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.