Voice Clone Tts
npc-live/clawfirm
声纹克隆和语音合成。上传音频样本克隆声纹,用克隆声纹或预设声纹生成语音。支持多个后端:MiniMax、ElevenLabs、Fish Audio、Azure TTS、OpenAI TTS。支持情绪控制、语速调整、批量生成。触发词:语音合成、TTS、声纹克隆、voice clone、text to speech、配音、旁白。
Turns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration.
SKILL.md written in Chinese; this summary is our English description.
$ npx skills add ConardLi/garden-skills --skill web-video-presentation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ConardLi/garden-skills web-video-presentation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ConardLi/garden-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/web-video-presentation .claude/skills/web-video-presentation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "web-video-presentation" agent skill from https://github.com/ConardLi/garden-skills/tree/main/skills/web-video-presentation into .claude/skills/web-video-presentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-video-presentation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ConardLi/garden-skills/tree/main/skills/web-video-presentationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ConardLi/garden-skills --skill web-video-presentation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ConardLi/garden-skills web-video-presentation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ConardLi/garden-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/web-video-presentation .agents/skills/web-video-presentation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "web-video-presentation" agent skill from https://github.com/ConardLi/garden-skills/tree/main/skills/web-video-presentation into .agents/skills/web-video-presentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-video-presentation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ConardLi/garden-skills --skill web-video-presentation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ConardLi/garden-skills web-video-presentation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ConardLi/garden-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/web-video-presentation .cursor/skills/web-video-presentation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "web-video-presentation" agent skill from https://github.com/ConardLi/garden-skills/tree/main/skills/web-video-presentation into .cursor/skills/web-video-presentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-video-presentation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ConardLi/garden-skills.git --path skills/web-video-presentation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ConardLi/garden-skills --skill web-video-presentation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ConardLi/garden-skills web-video-presentation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ConardLi/garden-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/web-video-presentation .gemini/skills/web-video-presentation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "web-video-presentation" agent skill from https://github.com/ConardLi/garden-skills/tree/main/skills/web-video-presentation into .gemini/skills/web-video-presentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-video-presentation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ConardLi/garden-skills web-video-presentationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ConardLi/garden-skills --skill web-video-presentation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ConardLi/garden-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/web-video-presentation .github/skills/web-video-presentation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "web-video-presentation" agent skill from https://github.com/ConardLi/garden-skills/tree/main/skills/web-video-presentation into .github/skills/web-video-presentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-video-presentation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ConardLi/garden-skills --skill web-video-presentation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ConardLi/garden-skills web-video-presentation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ConardLi/garden-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/web-video-presentation .opencode/skills/web-video-presentation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "web-video-presentation" agent skill from https://github.com/ConardLi/garden-skills/tree/main/skills/web-video-presentation into .opencode/skills/web-video-presentation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "web-video-presentation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
web-video-presentationTurns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration.
The agent starts from your article or script and produces a narration script and an outline in one pass, then asks you to settle five things together: the script, the outline, a theme, your assets and how development should run (chapter by chapter, in sequence or in parallel). The outline fixes pacing and information density only; animation is designed per chapter while it is built.
The output is a Vite and React project in which each click advances one beat of the narration and every step fills the screen, with a progress bar that only shows on hover. A narrations.ts file is the single source for the step count and the audio, so the script, outline, chapter code, chapter list and audio files cannot drift apart.
Audio synthesis is optional and not tied to one vendor: MiniMax mmx-cli and OpenAI TTS are built in, and ElevenLabs, edge-tts, Azure or your own TTS can be swapped in. The script, outline and each finished chapter go through a checklist review, preferably by a separate reviewer agent, and failures are fixed before the agent reports back. The instructions are written in Chinese.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit aaf9a82. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
npmbashnpxFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm and npx, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Web Video Presentation loads about 3.5k tokens when it runs, and up to ~32k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 941 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from ConardLi/garden-skills at commit aaf9a82, republished under its MIT licence (© ConardLi). 941 words, ~3,509 tokens.
.claude/skills/web-video-presentation/SKILL.md (or your agent's skills folder). This skill also uses 129 other files; get the full folder from GitHub.把一篇文章或口播稿,一步步做成可录屏的"伪装成视频的网页",可选合成 口播音频。产出物 = Vite + React + TS 项目 + 按章节切分的音频。
本 Skill 以方法论 + 协作流程为核心。脚手架模板提供 token 和原语, 但每个美学决策(配色、字型、动效气质)都应该针对你的主题重新设计 —— 不要照搬。
Phase 1 内容编写
1.1 识别用户输入
1.2 一次产出 script.md + outline.md
(口播稿 + 开发计划)
▼
[Checkpoint Plan] ← 必须停。一次对齐 5 件事:
稿子 / outline / 主题 / 素材 / 开发模式
▼
Phase 2 网页开发
2.1 脚手架(按选定主题)
2.2 第 1 章 = 主线程 + 完整版本(强制 anchor)
▼
[硬节点] 用户验收第 1 章 ← 不可跳过
▼
2.3 第 2~N 章(按选定模式:A 逐章 / B 顺序 / C 并行)
▼
[Checkpoint Audio] ← 必须停。是否合成音频
▼
Phase 3 音频合成(可选)
▼
Phase 4 录屏 + 后期工作目录约定(agent 在用户当前目录下创建 / 编辑):
my-video/
├── article.md # 用户给原文时必有 —— 不删!开发阶段画面信息源
├── script.md # 必有:保持原文语言的平台化口播稿(决定节拍)
├── outline.md # 必有:开发计划(章节切分 + 每步内容 + 信息池)
└── presentation/ # 脚手架产出的 Vite + React + TS 项目
├── src/chapters/<NN>-<id>/
│ ├── <Chapter>.tsx # 视觉实现
│ ├── <Chapter>.css
│ └── narrations.ts # ★ step 数 + 口播文本的唯一真相源
├── scripts/
│ ├── extract-narrations.ts # 扫所有 narrations.ts → audio-segments.json
│ ├── synthesize-audio.sh # provider-agnostic runner(循环 segments)
│ └── tts-providers/ # 每 provider 一个 .sh(内置 2 个)
│ ├── README.md # 三函数契约 + 5 段现成代码片段(11labs / edge-tts / say / azure / gcloud)
│ ├── minimax.sh # 默认 provider,用 mmx-cli
│ └── openai.sh # 内置 OpenAI TTS(curl + OPENAI_API_KEY)
├── audio-segments.json # extract 产出(合成前 review)
└── public/audio/<id>/<N>.mp3 # 可选:合成的音频关键:
narrations.ts是 step 数和音频合成的唯一真相源。 章节.tsx里的if (step === N)出现的最大 N + 1 必须等于narrations.length。这保证 5 处地方(script / outline / 章节代码 / chapters.ts / 音频文件)永远不会漂。
下面三个产出,每一个完成后必须走自检 → 修复 → 再汇报 / 推进:
| 产出 | 自检清单出处 |
|---|---|
script.md | SCRIPT-STYLE.md 三层自检(形式 / 风骨 / 念出来) |
outline.md | OUTLINE-FORMAT.md 自检 |
| 单章实现完成 | CHAPTER-CRAFT.md 完工自检 |
执行方式(按能力降级,优先用更隔离的方式):
铁律:拿到结论后先按 fail 项把产出改完,再向用户汇报"做完了
不同阶段读不同的文件。长会话里 agent 容易遗忘原则,特别是 Phase 2.4 的"实现单章"会重复 N 次 —— 每次都要回看核心约束。
| 阶段 | 必读(每次都看) | 一次性看完 / 按需查 |
|---|---|---|
| Phase 1.1-1.2 内容编写 | references/SCRIPT-STYLE.md + references/OUTLINE-FORMAT.md + article.md(用户原文,如有) | —— |
| Checkpoint Plan 选主题 | —— | themes/*/theme.json(动态读全部,列清单 + bestFor 推荐 + descriptionZh);references/THEMES.md(用户想了解主题系统时) |
| Phase 2.1 脚手架 | —— | SKILL.md 本节看一次 |
| Phase 2.4 实现单章(×N 次,被 2.2 / 2.3 调用) | references/CHAPTER-CRAFT.md 单一入口 —— Part 0 十条原则 / Part 1 开工 5 问 / Part 2 关系→动作决策树 / Part 3 视觉工具箱 / Part 4 时长参考 / Part 5 反 AI 味反模式 / Part 6 代码硬规则(含 narrations.ts 强制约束)/ Part 7 完工自检 / Part 8 反馈速查 + 当前主题的 themes/<id>/theme.json + 当前章节的 outline.md 段落 + article.md 本章对应段落 + 素材清单 | references/EXAMPLES/(结构示意,不是抄袭模板);references/THEMES.md 完整 token 契约 |
| Phase 3 音频合成 | references/AUDIO.md(含 narrations.ts → segments.json → 任意 provider 流程,内置 minimax + openai) | templates/scripts/tts-providers/README.md(换 provider / 自带 TTS 时) |
| Phase 4 录屏 + 后期 | references/RECORDING.md(含 ?auto=1 自动录屏) | —— |
| 选 / 造 / 切主题 | —— | references/THEMES.md |
写章节时只读一份
CHAPTER-CRAFT.md。十条原则 / 开工 self-prompting / 决策树 / 反 AI 味反模式 / 完工自检全部并入这一份单一入口。EXAMPLES/不是必读 —— 先按内容自由设计,卡壳才翻(按 anchor 翻"形",不要照搬)。
| 用户给的东西 | 该做的 |
|---|---|
| 原始文章(书面语 / 公众号 / 论文 / 博客) | 一次产出 script.md + outline.md(1.2),过 Checkpoint Plan |
| 直接的口播稿 / 视频脚本 | 落盘成 script.md,一次产出 outline.md(1.2 简化版),过 Checkpoint Plan |
| 啥都没有,只说"帮我做个 X 主题的视频" | 反问:先给一段素材或大纲。Skill 不替用户构思内容 |
两份产出物在一次思考中完成:
script.md:按 references/SCRIPT-STYLE.md
的规则把 article 转成保持原文语言的平台化口播稿。保留 article.md 不删——它是
outline 写信息池和章节实现画面时的细节源(双源原则)。outline.md:按 references/OUTLINE-FORMAT.md
规则切章节 + 切 step + 每章首段抽信息池。outline 的边界(关键):
| outline 必须写 | outline 不要写 |
|---|---|
| 章节切分 / 每章 step 数 / 估时 | 具体动画类型(blur clear / wipe / 弹簧) |
| 每步屏幕内容(hero / 数据 / 标语 / 列表项) | CSS 实现手段(filter / SVG / clip-path) |
| 章节级信息池:从 article 抽的数字 / 引用 / 案例 / 标签 | 时长数值(不写 |
| 步级关系名前缀("反差对照" / "递进列表" / "金句" 等可选 hint) | 持续微动 / 错峰量等微观节奏 |
outline 不写动画的理由:写死动画 = chapter agent 退化为翻译机; 留白让 chapter agent 在每步开工时按
CHAPTER-CRAFT.md的"内容驱动决策树"自由设计,才有真正的视频感。详见CHAPTER-CRAFT.mdPart 0 原则 7。
落盘后必须先走自检再进 Checkpoint Plan:按上文「硬性自检协议」分别
对 script.md / outline.md 执行(优先 Agent Teams → subAgent → 自检),
按结论修复完成后再进入 Checkpoint Plan。
script.md + outline.md 写完后必须停下来。用户在这一个节点同时确认
5 件事。
themes/*/theme.json 拿 nameZh / descriptionZh / bestFor
/ mood —— 不要硬编码清单script.md 的内容类型 / 关键词 / 语气,主动从主题里挑 2~3
套最匹配的推荐(匹配 bestFor 字段)outline.md 末尾"素材清单"部分内容计划写完,产出文件:
📄 article.md {若用户给原文则保留}
📄 script.md {X} 字 / ~{T} 分钟
📄 outline.md {N} 章 / {M} 步 + 每章信息池 + 末尾素材清单
章节速览:
1. <id> <章节标题> <S> 步 ~<T>s
2. ...
接下来一次对齐 5 件事:
1. 稿子 (script.md) 要不要改?
可以直接编辑文件,或口头告诉我修改方向。
2. 开发计划 (outline.md) 要不要改?重点看:
- 章节切分 / step 数 / 估时是否合理(合理判断:每章 30~60s)
- 每步屏幕内容是否清晰
- 每章首段「信息池」是否有足够的 article 细节供画面挂
- 末尾素材清单是否完整
3. 选哪个主题?我的推荐:
★ <推荐 1:nameZh (id)> — 因为 <bestFor 命中>;<descriptionZh 摘要>
★ <推荐 2 / 推荐 3>
其它可选:<剩余主题,nameZh + 一句话>
也可以让我帮你做新主题(详见 references/THEMES.md)。
4. 真素材怎么准备?粗看本视频要的图:<列粗略清单>
a) 我从 <现有素材路径> 帮你挑 b) 你自己提供 c) 全部 placeholder
5. 开发模式选哪个?
**第 1 章无论哪种模式都必须主线程做完 + 用户验收**(强制 anchor)。
差异在第 2 章及之后:
A) 默认 · 逐章确认(推荐)
每章做完都暂停验收 → 风险可控 / 节奏最稳
B) 第 1 章后顺序开发(不并行)
第 2~N 章主线程顺序做完后统一验收 → 速度中 / 适合 agent 不支持并行
C) 第 1 章后并行开发(subagent)
第 2~N 章用 subagent 并行 → 最快 / 用户控并行数(一次几章)
⚠️ 风格各章会有差异(这是预期,主题禁区兜底)收到反馈后:
bash <path-to-web-video-presentation>/scripts/scaffold.sh \
./presentation \
--theme=<用户选的主题 id>
bash <path-to-web-video-presentation>/scripts/scaffold.sh --list-themes自定义主题 → 先按
references/THEMES.md"创作新主题"流程做一个themes/<my-theme>/,再--theme=<my-theme>。
脚手架带一个 01-example demo。在写第一章真实内容前删掉:
rm -rf presentation/src/chapters/01-example并把 presentation/src/registry/chapters.ts 里 EXAMPLE_CHAPTER
的 import 和数组项移除。
核心:第 1 章 = 完整版本一次到位(节奏 + 视觉 + 真素材齐全)。 没有"骨架版"概念 —— 第一章就要做出用户能直接验收的样板。
为什么第 1 章必须主线程:
CHAPTER-CRAFT.md 这套指引在当前
主题 + 当前题材下的第一次落地做完第 1 章后必须停下来等用户验收:
第 1 章 <id> 做完了,dev server 在 localhost:5173 运行。
验收重点:
□ 视觉气质对不对?符合 <theme nameZh> 的预期吗?
□ 节奏对不对?某些步太快 / 太慢 / 信息太薄?
□ 内容驱动动画是否到位?还是有几步是无脑入场动画?
□ 双源原则:屏幕画面有没有"口播没念但 article 能挂"的细节?
□ 反 AI 味检查:紫粉渐变 / 圆角彩色边框 / 假插画 / emoji 是否有?
问题告诉我,我针对性改。OK 了告诉我"继续",我按选定模式做第 2 章及之后。所有模式下的共同规则:每章独立按 CHAPTER-CRAFT.md
开发。风格不强求章节间完全一致 —— 主题颜色 / 字体 token 兜底视觉
统一,动画 / 节奏 / 视觉演示由章节自由发挥是设计预期。
第 2 章做完 → 暂停验收 → OK → 第 3 章 → 暂停 → ... → 第 N 章。每章 独立验收,问题随时改,风险最低,节奏最稳。用户不明确选模式时 默认走这个。
第 2 章 → 第 3 章 → ... → 第 N 章 主线程顺序做完,最后统一验收。 速度中等,适合 agent 不支持并行任务的环境。
用 subagent 把第 2~N 章并行做完,最大并行数由用户控制("一次 4 章" / "一次 2 章")。最快,但风格各章会有差异 —— 这是预期,因为:
并行 subagent 的 prompt 必须包含:
references/CHAPTER-CRAFT.md 的路径(单一必读 —— 视觉演示要求 +
逐步揭示 + 双源原则 + 反 AI 味 + 代码红线 + 完工自检全部在这一份里)theme.json 的 descriptionZh / mood / bestFor(参考气质
即可,动画 / 时长 / 字号 / emoji 由 chapter agent 自由决定).cd- / .mg- / .pm- / ...);
不修改 chapters.ts;完工跑 npx tsc --noEmit重要:无论选哪种模式,用户随时可以中途切换模式。第 2 章 OK 后用户说"剩下的并行" / "剩下的逐章" 都行。
详细指引见 references/CHAPTER-CRAFT.md ——
单一必读入口,覆盖:视觉演示要求 / 逐步揭示 / 内容取舍 / 双源原则
/ 视频演示基本审美 / 反 AI 味 / 代码红线 / 完工自检。
核心要点(CHAPTER-CRAFT.md 详述):
改动 chapters.ts(增加 / 删除 / 重排章节,或某章 narrations.ts
长度变化)后,bump presentation/src/hooks/useStepper.ts 的
STORAGE_KEY(如 v4 → v5),避免持久化游标落到不存在的 step 上。
Phase 2 结束后必须停下来,问用户:
网页做完,{N} 章 {M} 步,dev server 在 localhost:5173 跑着。
要不要合成音频做"自动播放录屏"?
✓ 合成 → 扫所有章节的 narrations.ts 出 audio-segments.json,
调 TTS provider 合成每步一个 mp3 到 public/audio/。
合成完后用 ?auto=1 模式可以一镜到底录屏(音视频天然同步)。
内置两个 provider:
• minimax (mmx-cli) —— 默认,中文音色稳
• openai (OPENAI_API_KEY) —— curl-based,多数已有 key
其它后端 (ElevenLabs / edge-tts 免费 / macOS say 离线 /
Azure / Google) 见 scripts/tts-providers/README.md 的现成片段。
✗ 不合成 → 跳过 Phase 3,直接 Phase 4 用手动录屏 + 后期配音。要合成 → Phase 3。不合成 → 直接 Phase 4。
详细流程见 references/AUDIO.md。简版:
cd presentation
npm run extract-narrations # 扫所有 narrations.ts → audio-segments.json
# 让用户扫一眼 audio-segments.json 确认文本对
npm run synthesize-audio # 默认 minimax provider,增量
# 或用内置 openai (要 OPENAI_API_KEY):
PRESENTATION_TTS=openai npm run synthesize-audio
# 或自定义:写一个 scripts/tts-providers/<name>.sh,见该目录的 README.md合成完告诉用户:输出位置 / 总段数 / 哪些段时长异常(太长 = 该 step 拆 分;太短 = 文案太薄)—— 给最后一次校准节奏的机会。然后进入 Phase 4。
详见 references/RECORDING.md。两种路径:
| 场景 | 推荐路径 |
|---|---|
| Phase 3 已合成音频 | Auto 模式一镜到底:浏览器开 localhost:5173/?auto=1 → 按 SPACE → 整片自动播完 → 停录 → 裁头尾即成片,无需后期对音轨 |
| Phase 3 跳过 | 默认 Manual 模式手动点击推进 → 后期任意剪辑工具配音 |
agent 在 Phase 3 / Checkpoint Audio 后主动告诉用户适合的录屏路径。
完整展开见 references/CHAPTER-CRAFT.md
Part 0 —— 写章节时回那里查,下面只是索引。
| # | 原则 | 一句话 |
|---|---|---|
| 1 | 16:9 固定舞台 | 内容 1920×1080 + transform scale,没有响应式 |
| 2 | 全局 step 计数器 | 章节是 step 的纯函数,无定时器 |
| 3 | 每步独占整屏 | if (step === N) return <FullScene /> |
| 4 | 口播节拍 = step | 一节拍 = 一 step = 一聚焦想法 |
| 5 | 隐藏的边角控件 | 进度条 / 翻页器默认 opacity 0 |
| 6 | 舞台无 chrome | 没有 header / footer / 页码 / 品牌条 |
| 7 | 内容驱动动画 | 先找内在动作,找不到才入场动画兜底;持续微动慎用 |
| 8 | 多点逐个揭示 | 1 项 = 1 step,禁同步 stagger 上 N 项 |
| 9 | 整片同一主题 | 章节间不翻表面色;颜色 / 字体走 token,其它尺度章节自由 |
| 10 | 双源原则 | script 定节拍,article 定画面密度(落到信息池) |
简化表见 references/CHAPTER-CRAFT.md
Part 8「常见反馈速查」。关键:先定位是哪一层(节奏 / 视觉 / 内容
/ 代码),再改最小切片,不要重做整章。
按"何时读"标注,避免一次性全读:
| 文件 | 何时读 | 内容 |
|---|---|---|
references/SCRIPT-STYLE.md | Phase 1.2 必读 | 文章 → 口播稿规则、平台变体 |
references/OUTLINE-FORMAT.md | Phase 1.2 必读 | outline.md 字段 spec、命名约定、章节切分、信息池 |
references/CHAPTER-CRAFT.md | Phase 2.4 每章单一必读入口 | Part 0 十条原则 / Part 1 开工 5 问 / Part 2 关系→动作决策树 / Part 3 视觉工具箱 / Part 4 时长 / Part 5 反 AI 味反模式 / Part 6 代码硬规则 / Part 7 完工自检 / Part 8 反馈速查 |
references/EXAMPLES/ | 可选 —— 看结构 | 章节结构示意(hook / list-reveal / case-tech-review);不是抄袭模板 |
references/THEMES.md | 选 / 造 / 切主题时 | 完整 token 契约 + 内置主题清单 + 创作流程 |
references/AUDIO.md | Phase 3 才读 | provider-agnostic 音频合成流程、内置 minimax 用法、换 provider 路径、故障排查 |
templates/scripts/tts-providers/README.md | 换 / 加 TTS provider 时 | 三函数契约 + 内置 2 个 (minimax / openai) + 5 种现成代码片段(ElevenLabs / edge-tts / macOS say / Azure / Google) |
references/RECORDING.md | Phase 4 才读 | 录屏工具 + 后期合成 |
themes/ | Checkpoint Plan / Phase 1.2 时翻 | 内置主题(每个含 theme.json + tokens.css) |
scripts/scaffold.sh | Phase 2.1 跑一次 | 一键项目脚手架 |
© ConardLi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 129 other files (scripts, references) in skills/web-video-presentation of ConardLi/garden-skills.
Open the folder on GitHubat commit aaf9a82
Web Video Presentation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Web Video Presentation this skillConardLi/garden-skills | 13k | — | ~3.5k | Automated safety check: Pass | MIT | |
| Voice Clone Ttsnpc-live/clawfirm | 156 | — | ~1.1k | Automated safety check: Pass | None | |
| KrillinAI CLI Operatorkrillinai/OpenCreator | 13k | — | ~869 | Automated safety check: Pass | Apache-2.0 | |
| Video Productionspeechlab0210/video-production-skill | 105 | — | ~4.1k | Automated safety check: Notes | MIT | |
| Super Video MakerBomx/super-video-maker-skill | 308 | — | ~11k | Automated safety check: Notes | None | |
| Motion Videobestagentkits/motion-video-skill | 113 | — | ~1.5k | Automated safety check: Pass | MIT |
npc-live/clawfirm
声纹克隆和语音合成。上传音频样本克隆声纹,用克隆声纹或预设声纹生成语音。支持多个后端:MiniMax、ElevenLabs、Fish Audio、Azure TTS、OpenAI TTS。支持情绪控制、语速调整、批量生成。触发词:语音合成、TTS、声纹克隆、voice clone、text to speech、配音、旁白。
krillinai/OpenCreator
Routes agents to the right KrillinAI command for subtitles, dubbing, video rendering, covers and speech, and explains how to read its JSON and manifest output.
speechlab0210/video-production-skill
AI educational video production pipeline. An agent skill from speechlab0210/video-production-skill.
Bomx/super-video-maker-skill
End-to-end AI video production skill for agentic frameworks.
bestagentkits/motion-video-skill
Produce beat-synced 1080p motion-graphic videos in HyperFrames (HTML + GSAP) with an AI voice-over, Vietnamese karaoke captions, SFX and generated music, in one of two proven styles (glass keynote…
gooseworks-ai/goose-skills
Render a 'hand-swipe flavor listicle' video from a config — the brand's REAL product cutouts each float on a flat flavor-matched color field under a persistent title and slide RIGHT-TO-LEFT one to…
ConardLi/garden-skills
Turns a URL, PDF, DOCX, Markdown file, text or screenshots into a designed, shareable single-file HTML article through a staged review workflow.
ConardLi/garden-skills
Generates or edits images with GPT Image 2 through an OpenAI-compatible API, or falls back to writing finished prompts from a library of structured templates.
ConardLi/garden-skills
Builds polished browser-rendered HTML, CSS, JavaScript and React pages, dashboards, prototypes and slide decks, aiming for a stunning rather than merely functional result.
Categories
Turns an article or spoken script into a click-through, full-screen 16:9 web presentation that looks like a video, with optional synthesized narration. The agent starts from your article or script and produces a narration script and an outline in one pass, then asks you to settle five things together: the script, the outline, a theme, your assets and how development should run (chapter by chapter, in sequence or in parallel). The outline fixes pacing and information density only; animation is designed per chapter while it is built.
Web Video Presentation fits situations like: turning a spoken script or article into an explainer you can screen-record for a video platform; building a dynamic talk or keynote demo that has a cinematic feel and does not look like slides; producing a narrated product walkthrough with one audio file per chapter; recording a teaching tutorial as a 16:9 page with large type and one idea per screen.
Run `npx skills add ConardLi/garden-skills --skill web-video-presentation -a claude-code`. Or copy the skill folder (skills/web-video-presentation in ConardLi/garden-skills) into .claude/skills/web-video-presentation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ConardLi/garden-skills --skill web-video-presentation -a codex`. Or copy the skill folder (skills/web-video-presentation in ConardLi/garden-skills) into .agents/skills/web-video-presentation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ConardLi/garden-skills --skill web-video-presentation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/web-video-presentation, .gemini/skills/web-video-presentation, .github/skills/web-video-presentation and .opencode/skills/web-video-presentation in your project.
Going by SKILL.md and its folder, Web Video Presentation needs the command-line tools its instructions call (npm, bash and npx) and credentials named OPENAI_API_KEY. Our summary lists: A source article or narration script; A text-to-speech provider such as MiniMax mmx-cli or OpenAI TTS, only if you want audio.
SKILL.md contains no URLs. Its commands use npm and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Web Video Presentation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 28k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Web Video Presentation: Voice Clone Tts (npc-live/clawfirm, 156 stars), KrillinAI CLI Operator (krillinai/OpenCreator, 13k stars), Video Production (speechlab0210/video-production-skill, 105 stars) and Super Video Maker (Bomx/super-video-maker-skill, 308 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ConardLi (a GitHub user) maintains it in ConardLi/garden-skills, which has 12,767 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on July 12, 2026.
Source: ConardLi/garden-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.