Agent skill

Video Voiceover

by zenstory-ai in zenstory-ai/video-recap-skills

把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。

MITAuto-check passedMedia & Creative

Install Video Voiceover

skills CLI
$ npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zenstory-ai/video-recap-skills video-voiceover --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zenstory-ai/video-recap-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-voiceover .claude/skills/video-voiceover && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-voiceover
GitHub stars
561
Token cost
~1.6k tokens
SKILL.md length
350 words
Files
11 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。

  • Works in 9 steps: 定位 → 远程服务与数据外发 → 环境要求 → …
  • Tasks that involve Text to speech and voice
  • SKILL.md covers 1. 定位, 2. 远程服务与数据外发, 3. 环境要求 and 4. 输入契约, plus 5 more sections
  • Runs Python scripts from its folder; calls python3; reaches api.fish.audio; needs FISH_API_KEY and MIMO_API_KEY

What it does

Video Voiceover is an agent skill from zenstory-ai/video-recap-skills. 把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。 外发说明:每段旁白文字会发给所选 TTS 服务(MiMo / Fish Audio / 用户自托管的 IndexTTS),--voice-ref 的参考音频会发给 MiMo。 另含实验性的英译中 dub 路径:只在显式选择 dub 模式并传 --confirm-voice-rights 时运行,会把源视频音频发到 MiMo ASR, 并以原说话人的声音为参考经 MiMo voiceclone 克隆配音;只可用于用户有权使用、且说话人同意被克隆声音的内容。 触发词:配音、语音合成、TTS、解说配音、 voiceover、text to speech、旁白配音。

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 13 other files, including scripts and reference files (for example `references/index-tts.md`, `scripts/approved_text_policy.py` and `scripts/dub.py`).

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: Claude Code / Codex skills that turn a video into a Chinese narration recap (视频解说): scene detection, ASR, VLM, script, TTS, ffmpeg assembly, optional editable JianYing / CapCut… The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “/video-voiceover”

Requirements

  • Python 3
  • A credential in MIMO_TTS_API_KEY
  • A credential in MIMO_API_KEY

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. 定位
  2. 远程服务与数据外发
  3. 环境要求
  4. 输入契约
  5. 运行命令
  6. 输出契约
  7. 运行规则
  8. 实验性 dub 配音(英译中、克隆原声)
  9. 能力边界

What it can do on your machine

Read from SKILL.md and the folder at commit 5391686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.fish.audio

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FISH_API_KEY
    • MIMO_API_KEY
    • MIMO_TTS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Voiceover loads about 1.6k tokens when it runs, and up to ~2k if it reads all its reference files. Until then it costs about 119 tokens; SKILL.md has 350 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from zenstory-ai/video-recap-skills at commit 5391686, republished under its MIT licence (© zenstory-ai). 350 words, ~1,605 tokens.

Download SKILL.mdSave it as .claude/skills/video-voiceover/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
video-voiceover
description
把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 tts_segments 与 tts_meta.json。 外发说明:每段旁白文字会发给所选 TTS 服务(MiMo / Fish Audio / 用户自托管的 IndexTTS),--voice-ref 的参考音频会发给 MiMo。 另含实验性的英译中 dub 路径:只在显式选择 dub 模式并传 --confirm-voice-rights 时运行,会把源视频音频发到 MiMo ASR, 并以原说话人的声音为参考经 MiMo voiceclone 克隆配音;只可用于用户有权使用、且说话人同意被克隆声音的内容。 触发词:配音、语音合成、TTS、解说配音、 voiceover、text to speech、旁白配音。
user-invocable
false

1. 定位

本技能读取带时间戳的旁白稿,为每一段生成独立音频,并把语音适配到对应时间窗,随后记录下游合成所需的放置元数据。 默认引擎是 MiMo TTS(mimo-v2.5-tts);也可显式选择 Fish Audio(默认模型 s2.1-pro-free)。

2. 远程服务与数据外发

本技能只在下表中被选中的路径上联网,每条路径发出的内容如下。实读文本指 narration.json 段落去掉格式与舞台提示后的文字。

路径服务与凭据发送内容
默认解说 TTS(--tts-provider auto|mimo-tts)MiMo chat 接口 <MIMO_TTS_API_URL 或 MIMO_API_URL>/chat/completions,MIMO_TTS_API_KEY 或 MIMO_API_KEY,模型 mimo-v2.5-tts每段实读文本、一句自然语言语气/语速指令(由段的 emotion 与时间窗算出)、内置音色名(默认 冰糖);auto 下 MiMo key 缺失且设置了 FISH_API_KEY 时改走 Fish Audio 行
解说声音克隆(--voice-ref <audio> / VOICE_REF)同上,模型 mimo-v2.5-tts-voiceclone上述内容,外加参考音频(转成 24 kHz 单声道 WAV,最长 30 秒)的 base64,每段请求都带
Fish Audio(--tts-provider fish-audio)FISH_TTS_API_URL(默认 https://api.fish.audio/v1/tts),FISH_API_KEY每段实读文本、数值语速、音色 ID FISH_TTS_REFERENCE_ID;不发送本地音频
自托管 IndexTTS(--tts-provider index-tts)用户自己部署、由 INDEX_TTS_ENDPOINT 指定的 HTTP(S) 服务{"voice": INDEX_TTS_VOICE, "text": 实读文本};不发送本地音频
实验性 dub(见 §8)MiMo ASR(MIMO_API_URL,MIMO_API_KEY,mimo-v2.5-asr)与 MiMo voiceclone(TTS 接口与凭据,mimo-v2.5-tts-voiceclone)源视频整条音轨按 6 秒分窗送 ASR;每句中文译文连同从源音频截取的约 10 秒原说话人声音(克隆参考)送 voiceclone

--voice-ref 与 dub 都会把一个真实人物的声音发给 MiMo 用于克隆:只在用户有权使用该音频、且声音主人同意被克隆时使用。

3. 环境要求

bash
export MIMO_API_KEY=***  # 也可使用仅供 TTS 的 MIMO_TTS_API_KEY

# 或改用 Fish Audio TTS
export TTS_PROVIDER=fish-audio
export FISH_API_KEY=***
export FISH_TTS_REFERENCE_ID=<voice-model-id>  # 可选;覆盖内置“娱乐扒妹”音色

# 或显式选择自托管 index-tts 端点,配置见 references/index-tts.md
export TTS_PROVIDER=index-tts

下面的 scripts/... 均相对于本技能目录。若执行器从仓库根目录启动,请给脚本路径加上本技能的绝对目录。

4. 输入契约

默认输入为 work_dir/narration.json。每段必须包含 start、end 与 narration,可选字段包括 pause_after_ms 和 overlaps_speech。时间统一表示音频最终放置的输出时间线秒数。

cut 流程先剪后配:narration.json 本身就是按剪后成片的输出时间写的,不存在另一份映射稿。

5. 运行命令

bash
python3 scripts/voiceover.py --work-dir <work_dir> --narration <narration.json> \
  [--tts-provider auto|mimo-tts|fish-audio|index-tts] \
  [--mimo-voice 冰糖 | --voice-ref <reference-audio>] \
  [--preserve-approved-text] [--allow-partial-tts]

单独运行且省略 --narration 时,默认读取 work_dir/narration.json;--narration 只用于指定其他路径的同格式稿件。

6. 输出契约

  • tts_segments/*.wav:每段旁白对应一个音频文件。
  • tts_meta.json:包含 segments、engine、voice(实际使用的 provider、模型、音色或参考音频)与 narration。每段记录 audio_path、时间、 pause_after_ms 和放置字段。
  • 干净运行写入 partial: false 与 failures: []。
  • 使用 --allow-partial-tts 跳过失败段时,写入 partial: true 和 failures: [{index,start,end,text,error}],让缺失语音保持可见。
  • --preserve-approved-text 是显式的批准稿保护策略。每段保留原始 authored_text 证据;TTS 实际读取的 spoken_text 只经过既有的格式/舞台提示清理。若完整语音超过时间窗及 累计语速预算,命令失败并报告段序号、原稿、实读文本、语音时长和窗口证据,不写成功的 tts_meta.json。严格模式下任何必需段失败(包括供应商失败)都不能被 --allow-partial-tts 降级为可交付的部分成功;异常记录标为 required: true 并带策略 ID。 仅含 [停顿] 等清理标记、清理后无实读文本的作者段也属于必需段错误。
Show full SKILL.md (188 more words)Show less

7. 运行规则

  • 分段音频按内容缓存在 tts_segments/cache/:键是实读文本、实际发给供应商的语气请求(MiMo 是那句自然语言指令,语速只在 ≥+6% 或 ≤-3% 时改变措辞;Fish Audio 是数值 speed;index-tts 没有段级控制)与 TTS 设置,不含段序号和时间窗;因此段位变化让名义语速从 +5% 变成 -2% 时,MiMo 不重新合成; 缓存 WAV 的 size/mtime_ns 变了即失效。narr_NNN.wav 是指向缓存的硬链接(不支持时为副本),tts_meta.json 照旧引用它。删掉、插入或挪动某段后,只重生成文本或发给供应商的请求变了的段(名义语速随首段、末两段的位置变化); 旧版的 narr_NNN.wav.cache.json 不再读取。
  • 批准稿保护策略属于缓存设置:严格模式往缓存键里加入策略与原稿,与默认策略(report-over-budget-v2,不进键)互不命中;旧版逐段缓存(含自动缩稿音频)不再读取;只有同一严格策略下、 spoken_text 完整匹配且 WAV 存在非空的缓存才可离线复用;复用时仍按当前时间窗检查,放不下照样失败。
  • 严格 CLI 在本轮合成前把旧 tts_meta.json 按时间戳归档至 tts_meta.history/,因此失败时 当前路径不会继续冒充本轮成功;成功元数据通过同目录临时文件原子替换。
  • auto 优先使用已配置的 MiMo,MiMo key 缺失且设置了 FISH_API_KEY 时使用 Fish Audio;需要可复现的 provider 选择时显式传 --tts-provider。
  • 自托管 index-tts 端点只能由 --tts-provider index-tts 或 TTS_PROVIDER=index-tts 显式选择,auto 永不兜底选择它。协议、请求体、receipt 语义与缓存失效规则见 references/index-tts.md。
  • Fish Audio 直接请求 WAV;默认使用“娱乐扒妹”音色(5653cea4ac83480aaf2bf45406556185),FISH_TTS_REFERENCE_ID 可覆盖。模型、音色 ID、API URL、归一化设置或按内容计算出的语速变化时会重新生成缓存(Fish 不接收音高和情绪,它们变了不重新生成)。当前免费模型无 SLA,受 Fair Use 和官方免费期限约束。
  • --voice-ref 仅用于 full/cut 解说克隆,切换到 mimo-v2.5-tts-voiceclone。仅在确需新合成时惰性规范化一次; 参考音频的路径、size/mtime_ns 或预处理版本变化会使旧缓存失效。仅在获得授权后使用,参考音频会发送到 MiMo。
  • dub voiceclone 原始 WAV 也会按模型、提示、台词和参考音频的 size/mtime_ns 缓存;匹配重跑不再重复请求或计费, dub_manifest.json 逐行记录 tts_cache=hit|miss。
  • 合成出的段音频比按 TTS_MIN_SPEECH_RATE(默认 2.5 字/秒,英文按每词 1.5 字)读完全文、再加停顿与首尾静音的上限还长时, 视为 TTS 幻读(读完原稿后又编出一段话),按失败重试,不缓存也不交给 assemble;重试用尽则该段失败,报错写明时长与上限, 最后一次被拒的音频留在 tts_segments/narr_NNN.rejected.wav 供试听。数字(半角/全角)逐个计 1 字,% 计 3 字(百分之)。 dub 的 voiceclone 台词走同一道检查与重试(被拒的留在 dub_tts/line_NNN_raw.rejected.wav)。 旧版本缓存下的这类 WAV 在重跑时不再复用,会重新合成。设为 0 关闭这道检查。
  • TTS_WORKERS、TTS_TIMEOUT、TTS_RETRIES、ALLOW_PARTIAL_TTS 用于调整并发、超时、重试与部分成功策略。

8. 实验性 dub 配音(英译中、克隆原声)

dub.py 是实验功能:把英文原声翻译成中文,并用原说话人的克隆音色整轨替换人声。它与上面的解说配音是两条独立路径,voiceover.py 从不调用它。

  • 触发:只在用户明确要求英译中配音时,由编排入口以 --edit-mode dub --confirm-voice-rights 运行;编排入口把确认参数原样转给 dub.py 的准备和渲染两个阶段。没有单独的手动阶段。
  • 确认门禁:dub.py 的两个阶段都必须带 --confirm-voice-rights,缺少时在抽取音频和发出任何请求之前退出,并说明会外发什么。编排入口在 dub 模式下同样拒绝缺少该参数的运行,在其他模式下拒绝该参数。
  • 外发内容:准备阶段把源视频音轨按 6 秒分窗发给 MiMo ASR(mimo-v2.5-asr)转写英文;渲染阶段把每句中文译文和从源音频第 2 秒起截取的约 10 秒原说话人声音(dub_reference.wav)一起发给 MiMo voiceclone(mimo-v2.5-tts-voiceclone)。
  • 权利与同意:只能用于用户有权使用的视频与音频,且被克隆声音的说话人已同意。Agent 必须先向用户确认这两点,用户确认后才能加 --confirm-voice-rights;无法确认时不要运行 dub,也不要用它冒充他人发言。
  • 确定性门禁:渲染阶段在语音克隆前写 dub_lint.json,空行、重叠或越界译文即中止,不发 voiceclone 请求。
  • 本地产物:dub_source.wav、dub_transcript.json、dub_brief.md、dub_reference.wav、Agent 写的 dub_script.json、dub_lint.json、dub_tts/、dub_manifest.json 与 dub_<name>.mp4,都只写在 work_dir。

9. 能力边界

  • 超窗时保留原稿并记录日志;assemble 有界提速放不下则在渲染前以 no_safe_fit 阻断。批准稿加 --preserve-approved-text,超窗即在 TTS 阶段失败。
  • 不混流、不压低原声、不渲染字幕。
  • 不分析视频,也不选择时间点;只为输入稿件中的既定分段配音。
  • Fish Audio 与 IndexTTS 路径都不接受本地 --voice-ref;前者用已创建的 FISH_TTS_REFERENCE_ID 选择音色。

© zenstory-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts, references) in skills/video-voiceover of zenstory-ai/video-recap-skills.

  • SKILL.md
  • references/index-tts.md
  • scripts/approved_text_policy.py
  • scripts/dub.py
  • scripts/lib.py
  • scripts/providers/__init__.py
  • scripts/providers/fish_audio.py
  • scripts/providers/index_tts.py
  • scripts/tts_audio.py
  • scripts/tts_cache.py
  • scripts/voiceover.py

Open the folder on GitHubat commit 5391686

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders. This page covers the copy in zenstory-ai/video-recap-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Video Voiceover next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Voiceover compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Voiceover this skillzenstory-ai/video-recap-skills561—~1.6kAutomated safety check: PassMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo130k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from zenstory-ai/video-recap-skills

  • Video Recap

    zenstory-ai/video-recap-skills

    从输入视频生成中文解说成片或原声剧情短片。用户提供 .mp4 / .mov / .mkv / .webm,并要求剪辑、添加旁白、 配音、总结、短剧/电视剧/电影/纪录片/科普解说时使用。负责编排 video- 技能链:视频理解 → Agent 制定故事与视听方案 → 剪辑 → 配音 → 合成。触发词:视频解说、视频旁白、生成解说、 视频 recap、video…

    561 GitHub stars~2.4k tokensUpdated 7 days ago
    Auto-check passed
  • Video Assemble

    zenstory-ai/video-recap-skills

    合成视频解说最终成片:把旁白音频铺到源视频上,按旁白窗口压低原声,生成 SRT / ASS 字幕并可烧录, 最后做响度标准化。作为最终合成阶段使用。输入源视频、ttsmeta.json 与旁白位置; 输出 recap 成片和字幕。触发词:视频合成、混音、字幕、压字幕、assemble video、mux、ducking、subtitles、成片。

    561 GitHub stars~1.7k tokensUpdated 7 days ago
    Auto-check passed
  • Video Cut

    zenstory-ai/video-recap-skills

    把长视频按 Agent 选择的原片区间剪成短片。作为两阶段创作流程中的剪辑环节,读取 clipplan.json 与源视频, 输出 editedsource.mp4;随后 Agent 按输出时间线写 narration.json。支持单视频与多视频(sources manifest)拼剪, 本工具不读取、不映射旁白。

    561 GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed
  • Video Reference

    zenstory-ai/video-recap-skills

    按需把一部成片拆成可复用的制作参考:测镜头节奏与响度,标注段落与音轨分工,把原片事实与可迁移方法分开, 导出不含原片人名台词的 productionreference.json 供下次制作参考。不在默认生产路径上。

    561 GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Video Script

    zenstory-ai/video-recap-skills

    对已完成分析的视频进行导演与剪辑策划,再写带时间戳的中文解说并校验;也处理已有短片的 宣发标题、花字修订和外部文案回填。普通策划输入 workdir 的 agentnarrationbrief.md 与 vlmanalysis.json;文案返修输入当前成片的工程与内容证据。策划输出 recapstoryplan.json、visualaudioboard.json、 可选…

    561 GitHub stars~2.4k tokensUpdated 7 days ago
    Auto-check passed
  • Video Understanding

    zenstory-ai/video-recap-skills

    把视频分析为结构化理解索引:场景检测、ASR 转写、逐场景 VLM 观察、静音窗口、融合时间线和写作 brief. An agent skill from zenstory-ai/video-recap-skills.

    561 GitHub stars~1.1k tokensUpdated 7 days ago
    Auto-check passed

Questions about Video Voiceover

What does Video Voiceover do?

把带时间戳的 narration.json 合成为中文解说音频。使用 MiMo TTS(mimo-v2.5-tts)或 Fish Audio(s2.1-pro-free)或显式配置的通用 IndexTTS HTTP 服务逐段生成语音, 按时间窗动态适配语速并处理响度;输入输出时间线上的旁白,产出 ttssegments 与 ttsmeta.json。. Video Voiceover is an agent skill from zenstory-ai/video-recap-skills.

When should I use Video Voiceover?

Video Voiceover fits situations like: tasks that involve Text to speech and voice.

How do I install Video Voiceover in Claude Code?

Run `npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a claude-code`. Or copy the skill folder (skills/video-voiceover in zenstory-ai/video-recap-skills) into .claude/skills/video-voiceover in your project. Claude Code loads it when a task matches its description.

How do I install Video Voiceover in Codex?

Run `npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a codex`. Or copy the skill folder (skills/video-voiceover in zenstory-ai/video-recap-skills) into .agents/skills/video-voiceover in your project. Codex loads it when a task matches its description.

Can I use Video Voiceover in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zenstory-ai/video-recap-skills --skill video-voiceover -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-voiceover, .gemini/skills/video-voiceover, .github/skills/video-voiceover and .opencode/skills/video-voiceover in your project.

What does Video Voiceover need to run?

Going by SKILL.md and its folder, Video Voiceover needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named FISH_API_KEY, MIMO_API_KEY and MIMO_TTS_API_KEY. Our summary lists: Python 3; A credential in MIMO_TTS_API_KEY; A credential in MIMO_API_KEY.

Does Video Voiceover access the network?

SKILL.md names 1 domain. In commands or code: api.fish.audio; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Video Voiceover safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Voiceover use?

Video Voiceover is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Voiceover use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 406 tokens, read only when the agent opens those files.

What are the alternatives to Video Voiceover?

Skills that share tags, products or a category with Video Voiceover: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Voiceover?

zenstory-ai (a GitHub organization) maintains it in zenstory-ai/video-recap-skills, which has 561 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 4, 2026.

Source: zenstory-ai/video-recap-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.