MoneyPrinterTurbo Video Generator
harry0703/MoneyPrinterTurbo
Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.
把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。
$ npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-voiceover --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-voiceover .claude/skills/video-voiceover && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-voiceover" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-voiceover into .claude/skills/video-voiceover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-voiceover", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-voiceoverType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-voiceover --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/video-voiceover .agents/skills/video-voiceover && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-voiceover" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-voiceover into .agents/skills/video-voiceover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-voiceover", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-voiceover --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/video-voiceover .cursor/skills/video-voiceover && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-voiceover" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-voiceover into .cursor/skills/video-voiceover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-voiceover", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/zenstory-ai/video-recap-skills.git --path skills/video-voiceover--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-voiceover --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/video-voiceover .gemini/skills/video-voiceover && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-voiceover" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-voiceover into .gemini/skills/video-voiceover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-voiceover", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install zenstory-ai/video-recap-skills video-voiceoverInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/video-voiceover .github/skills/video-voiceover && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-voiceover" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-voiceover into .github/skills/video-voiceover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-voiceover", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install zenstory-ai/video-recap-skills video-voiceover --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/video-voiceover .opencode/skills/video-voiceover && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-voiceover" agent skill from https://github.com/zenstory-ai/video-recap-skills/tree/main/skills/video-voiceover into .opencode/skills/video-voiceover/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-voiceover", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-voiceover把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。
Video Voiceover is an agent skill from zenstory-ai/video-recap-skills. 把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。 外发说明:每段旁白文字会发给所选 TTS 服务(MiMo / Fish Audio / 用户自托管的 IndexTTS),--voice-ref 的参考音频会发给 MiMo。 另含实验性的英译中 dub 路径:只在显式选择 dub 模式并传 --confirm-voice-rights 时运行,会把源视频音频发到 MiMo ASR, 并以原说话人的声音为参考经 MiMo voiceclone 克隆配音;只可用于用户有权使用、且说话人同意被克隆声音的内容。 触发词:配音、语音合成、TTS、解说配音、 voiceover、text to speech、旁白配音。
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `references/index-tts.md`, `scripts/approved_text_policy.py` and `scripts/dub.py`).
It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: Claude Code / Codex skills that turn a video into a Chinese narration recap (视频解说): scene detection, ASR, VLM, script, TTS, ffmpeg assembly, optional editable JianYing / CapCut… The licence is MIT.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 5391686. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 9 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.fish.audioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
FISH_API_KEYMIMO_API_KEYMIMO_TTS_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Video Voiceover loads about 1.6k tokens when it runs, and up to ~2k if it reads all its reference files. Until then it costs about 119 tokens; SKILL.md has 350 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from zenstory-ai/video-recap-skills at commit 5391686, republished under its MIT licence (© zenstory-ai). 350 words, ~1,605 tokens.
.claude/skills/video-voiceover/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.本技能读取带时间戳的旁白稿,为每一段生成独立音频,并把语音适配到对应时间窗,随后记录下游合成所需的放置元数据。
默认引擎是 MiMo TTS(mimo-v2.5-tts);也可显式选择 Fish Audio(默认模型 s2.1-pro-free)。
本技能只在下表中被选中的路径上联网,每条路径发出的内容如下。实读文本指 narration.json 段落去掉格式与舞台提示后的文字。
| 路径 | 服务与凭据 | 发送内容 |
|---|---|---|
默认解说 TTS(--tts-provider auto|mimo-tts) | MiMo chat 接口 <MIMO_TTS_API_URL 或 MIMO_API_URL>/chat/completions,MIMO_TTS_API_KEY 或 MIMO_API_KEY,模型 mimo-v2.5-tts | 每段实读文本、一句自然语言语气/语速指令(由段的 emotion 与时间窗算出)、内置音色名(默认 冰糖);auto 下 MiMo key 缺失且设置了 FISH_API_KEY 时改走 Fish Audio 行 |
解说声音克隆(--voice-ref <audio> / VOICE_REF) | 同上,模型 mimo-v2.5-tts-voiceclone | 上述内容,外加参考音频(转成 24 kHz 单声道 WAV,最长 30 秒)的 base64,每段请求都带 |
Fish Audio(--tts-provider fish-audio) | FISH_TTS_API_URL(默认 https://api.fish.audio/v1/tts),FISH_API_KEY | 每段实读文本、数值语速、音色 ID FISH_TTS_REFERENCE_ID;不发送本地音频 |
自托管 IndexTTS(--tts-provider index-tts) | 用户自己部署、由 INDEX_TTS_ENDPOINT 指定的 HTTP(S) 服务 | {"voice": INDEX_TTS_VOICE, "text": 实读文本};不发送本地音频 |
| 实验性 dub(见 §8) | MiMo ASR(MIMO_API_URL,MIMO_API_KEY,mimo-v2.5-asr)与 MiMo voiceclone(TTS 接口与凭据,mimo-v2.5-tts-voiceclone) | 源视频整条音轨按 6 秒分窗送 ASR;每句中文译文连同从源音频截取的约 10 秒原说话人声音(克隆参考)送 voiceclone |
--voice-ref 与 dub 都会把一个真实人物的声音发给 MiMo 用于克隆:只在用户有权使用该音频、且声音主人同意被克隆时使用。
export MIMO_API_KEY=*** # 也可使用仅供 TTS 的 MIMO_TTS_API_KEY
# 或改用 Fish Audio TTS
export TTS_PROVIDER=fish-audio
export FISH_API_KEY=***
export FISH_TTS_REFERENCE_ID=<voice-model-id> # 可选;覆盖内置“娱乐扒妹”音色
# 或显式选择自托管 index-tts 端点,配置见 references/index-tts.md
export TTS_PROVIDER=index-tts下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。
默认输入为 work_dir/narration.json。每段必须包含 start、end 与 narration,可选字段包括
pause_after_ms 和 overlaps_speech。时间统一表示音频最终放置的输出时间线秒数。
cut 流程先剪后配:narration.json 本身就是按剪后成片的输出时间写的,不存在另一份映射稿。
python3 scripts/voiceover.py --work-dir <work_dir> --narration <narration.json> \
[--tts-provider auto|mimo-tts|fish-audio|index-tts] \
[--mimo-voice 冰糖 | --voice-ref <reference-audio>] \
[--preserve-approved-text] [--allow-partial-tts]单独运行且省略 --narration 时,默认读取 work_dir/narration.json;--narration 只用于指定其他路径的同格式稿件。
tts_segments/*.wav:每段旁白对应一个音频文件。tts_meta.json:包含 segments、engine、voice(实际使用的 provider、模型、音色或参考音频)与 narration。每段记录 audio_path、时间、
pause_after_ms 和放置字段。partial: false 与 failures: []。--allow-partial-tts 跳过失败段时,写入 partial: true 和
failures: [{index,start,end,text,error}],让缺失语音保持可见。--preserve-approved-text 是显式的批准稿保护策略。每段保留原始 authored_text
证据;TTS 实际读取的 spoken_text 只经过既有的格式/舞台提示清理。若完整语音超过时间窗及
累计语速预算,命令失败并报告段序号、原稿、实读文本、语音时长和窗口证据,不写成功的
tts_meta.json。严格模式下任何必需段失败(包括供应商失败)都不能被
--allow-partial-tts 降级为可交付的部分成功;异常记录标为 required: true 并带策略 ID。
仅含 [停顿] 等清理标记、清理后无实读文本的作者段也属于必需段错误。tts_segments/cache/:键是实读文本、实际发给供应商的语气请求(MiMo 是那句自然语言指令,语速只在 ≥+6% 或 ≤-3% 时改变措辞;Fish Audio 是数值 speed;index-tts 没有段级控制)与 TTS 设置,不含段序号和时间窗;因此段位变化让名义语速从 +5% 变成 -2% 时,MiMo 不重新合成;
缓存 WAV 的 size/mtime_ns 变了即失效。narr_NNN.wav 是指向缓存的硬链接(不支持时为副本),tts_meta.json
照旧引用它。删掉、插入或挪动某段后,只重生成文本或发给供应商的请求变了的段(名义语速随首段、末两段的位置变化);
旧版的 narr_NNN.wav.cache.json 不再读取。report-over-budget-v2,不进键)互不命中;旧版逐段缓存(含自动缩稿音频)不再读取;只有同一严格策略下、
spoken_text 完整匹配且 WAV 存在非空的缓存才可离线复用;复用时仍按当前时间窗检查,放不下照样失败。tts_meta.json 按时间戳归档至 tts_meta.history/,因此失败时
当前路径不会继续冒充本轮成功;成功元数据通过同目录临时文件原子替换。auto 优先使用已配置的 MiMo,MiMo key 缺失且设置了 FISH_API_KEY 时使用 Fish Audio;需要可复现的 provider 选择时显式传 --tts-provider。--tts-provider index-tts 或 TTS_PROVIDER=index-tts 显式选择,auto
永不兜底选择它。协议、请求体、receipt 语义与缓存失效规则见 references/index-tts.md。5653cea4ac83480aaf2bf45406556185),FISH_TTS_REFERENCE_ID 可覆盖。模型、音色 ID、API URL、归一化设置或按内容计算出的语速变化时会重新生成缓存(Fish 不接收音高和情绪,它们变了不重新生成)。当前免费模型无 SLA,受 Fair Use 和官方免费期限约束。--voice-ref 仅用于 full/cut 解说克隆,切换到 mimo-v2.5-tts-voiceclone。仅在确需新合成时惰性规范化一次;
参考音频的路径、size/mtime_ns 或预处理版本变化会使旧缓存失效。仅在获得授权后使用,参考音频会发送到 MiMo。size/mtime_ns 缓存;匹配重跑不再重复请求或计费,
dub_manifest.json 逐行记录 tts_cache=hit|miss。TTS_MIN_SPEECH_RATE(默认 2.5 字/秒,英文按每词 1.5 字)读完全文、再加停顿与首尾静音的上限还长时,
视为 TTS 幻读(读完原稿后又编出一段话),按失败重试,不缓存也不交给 assemble;重试用尽则该段失败,报错写明时长与上限,
最后一次被拒的音频留在 tts_segments/narr_NNN.rejected.wav 供试听。数字(半角/全角)逐个计 1 字,% 计 3 字(百分之)。
dub 的 voiceclone 台词走同一道检查与重试(被拒的留在 dub_tts/line_NNN_raw.rejected.wav)。
旧版本缓存下的这类 WAV 在重跑时不再复用,会重新合成。设为 0 关闭这道检查。TTS_WORKERS、TTS_TIMEOUT、TTS_RETRIES、ALLOW_PARTIAL_TTS 用于调整并发、超时、重试与部分成功策略。dub.py 是实验功能:把英文原声翻译成中文,并用原说话人的克隆音色整轨替换人声。它与上面的解说配音是两条独立路径,voiceover.py 从不调用它。
--edit-mode dub --confirm-voice-rights 运行;编排入口把确认参数原样转给 dub.py 的准备和渲染两个阶段。没有单独的手动阶段。dub.py 的两个阶段都必须带 --confirm-voice-rights,缺少时在抽取音频和发出任何请求之前退出,并说明会外发什么。编排入口在 dub 模式下同样拒绝缺少该参数的运行,在其他模式下拒绝该参数。mimo-v2.5-asr)转写英文;渲染阶段把每句中文译文和从源音频第 2 秒起截取的约 10 秒原说话人声音(dub_reference.wav)一起发给 MiMo voiceclone(mimo-v2.5-tts-voiceclone)。--confirm-voice-rights;无法确认时不要运行 dub,也不要用它冒充他人发言。dub_lint.json,空行、重叠或越界译文即中止,不发 voiceclone 请求。dub_source.wav、dub_transcript.json、dub_brief.md、dub_reference.wav、Agent 写的 dub_script.json、dub_lint.json、dub_tts/、dub_manifest.json 与 dub_<name>.mp4,都只写在 work_dir。no_safe_fit 阻断。批准稿加 --preserve-approved-text,超窗即在 TTS 阶段失败。--voice-ref;前者用已创建的 FISH_TTS_REFERENCE_ID 选择音色。© zenstory-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (scripts, references) in skills/video-voiceover of zenstory-ai/video-recap-skills.
Open the folder on GitHubat commit 5391686
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in zenstory-ai/video-recap-skills, which our catalogue first saw on October 7, 2026.
Video Voiceover next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Video Voiceover this skillzenstory-ai/video-recap-skills | 561 | — | ~1.6k | Automated safety check: Pass | MIT | |
| MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo | 130k | — | ~2.1k | Automated safety check: Warn | MIT | |
| HyperFrames Media Useheygen-com/hyperframes | 60k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Openspec OnboardSAP/e-mobility-charging-stations-simulator | 227 | 25 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Blog AudioAgriciDaniel/claude-blog | 2.3k | 1 repos | ~2.2k | Automated safety check: Notes | MIT | |
| Musictadaspetra/loop | 296 | 2 repos | ~827 | Automated safety check: Pass | MIT |
harry0703/MoneyPrinterTurbo
Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.
heygen-com/hyperframes
Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.
SAP/e-mobility-charging-stations-simulator
Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.
AgriciDaniel/claude-blog
Generate audio narration of blog posts using Google Gemini TTS.
tadaspetra/loop
Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.
hoquanghai/Auto-Create-Video
Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.
zenstory-ai/video-recap-skills
从输入视频生成中文解说成片或原声剧情短片。用户提供 .mp4 / .mov / .mkv / .webm,并要求剪辑、添加旁白、 配音、总结、短剧/电视剧/电影/纪录片/科普解说时使用。负责编排 video- 技能链:视频理解 → Agent 制定故事与视听方案 → 剪辑 → 配音 → 合成。触发词:视频解说、视频旁白、生成解说、 视频 recap、video…
zenstory-ai/video-recap-skills
合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。
zenstory-ai/video-recap-skills
把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。
zenstory-ai/video-recap-skills
按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。
zenstory-ai/video-recap-skills
对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验;也处理已有短片的 宣发标题、花字修订和外部文案回填。普通策划输入 workdir 的 agentnarrationbrief.md 与 vlmanalysis.json;文案返修输入当前成片的工程与内容证据。策划输出 recapstoryplan.json、visualaudioboard.json、 可选…
zenstory-ai/video-recap-skills
把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.
Categories
把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。. Video Voiceover is an agent skill from zenstory-ai/video-recap-skills.
Video Voiceover fits situations like: tasks that involve Text to speech and voice.
Run `npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a claude-code`. Or copy the skill folder (skills/video-voiceover in zenstory-ai/video-recap-skills) into .claude/skills/video-voiceover in your project. Claude Code loads it when a task matches its description.
Run `npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a codex`. Or copy the skill folder (skills/video-voiceover in zenstory-ai/video-recap-skills) into .agents/skills/video-voiceover in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-voiceover, .gemini/skills/video-voiceover, .github/skills/video-voiceover and .opencode/skills/video-voiceover in your project.
Going by SKILL.md and its folder, Video Voiceover needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named FISH_API_KEY, MIMO_API_KEY and MIMO_TTS_API_KEY. Our summary lists: Python 3; A credential in MIMO_TTS_API_KEY; A credential in MIMO_API_KEY.
SKILL.md names 1 domain. In commands or code: api.fish.audio; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Video Voiceover is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 406 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Video Voiceover: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
zenstory-ai (a GitHub organization) maintains it in zenstory-ai/video-recap-skills, which has 561 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 4, 2026.
Source: zenstory-ai/video-recap-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.