Agent skill

AI Video Production Assistant

by wanghui2323 in wanghui2323/ai-video-maker

Turns an idea, article, outline or audio file into a sourced, reviewable AI video, tracking whether narration uses a human, synthetic or cloned voice.

MITAuto-check passedMedia & Creative

SKILL.md written in Chinese; this summary is our English description.

Install AI Video Production Assistant

skills CLI
$ npx skills add wanghui2323/ai-video-maker --skill make-ai-video -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wanghui2323/ai-video-maker make-ai-video --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wanghui2323/ai-video-maker.git skills-src && mkdir -p .claude/skills && cp -r skills-src/make-ai-video .claude/skills/make-ai-video && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
make-ai-video
GitHub stars
101
Token cost
~924 tokens
SKILL.md length
177 words
Files
32 (incl. scripts, references, assets)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Turns an idea, article, outline or audio file into a sourced, reviewable AI video, tracking whether narration uses a human, synthetic or cloned voice.

  • Works in 9 steps: 建立 video-brief.json → 检查声音依赖 → 建立 content-decision.json → …
  • Turning a written idea or article into a planned video with narration and visuals
  • SKILL.md covers 启动即执行, 读取所需合同, 使用两条时钟 and 执行工作流, plus 2 more sections
  • Calls node

What it does

Runs a diagnostic script on activation to check what already exists in the current production package, then writes whatever the user has provided into a video brief and keeps moving forward until it reaches the next real approval point that needs a human decision, rather than just listing steps or handing commands back to the user. Reversible local actions such as checking files, creating folders, and validating the brief proceed on their own; installing dependencies or downloading a large voice model requires explaining size and purpose first and getting approval.

Two separate cycles govern the work: an infrequent voice-setup cycle that only runs when the user explicitly wants a cloned voice and no usable voice profile exists yet, covering consent, reference recordings and calibration candidates; and a per-video production cycle that goes from the brief through a content decision, a content plan, locked narration, generated audio, and on to rendering. Content decisions are generated as several candidates that each answer one question, with only one recommended at a time while the rest are kept.

Video length follows fixed bands by content density, from a 35-to-50-second single correction up to a multi-section deep piece, and full rendering or cloned narration is blocked until both the content plan and the narration have been explicitly reviewed. A bundled example package shows the complete set of objects for reference, and the skill is told never to invent personal experience, data or examples to fill a gap.

When your agent uses it

  • Turning a written idea or article into a planned video with narration and visuals
  • Checking current production status before resuming an in-progress video
  • Setting up a cloned voice profile before using it in a video
  • Reviewing a locked content plan and narration before rendering starts

Example prompts

  • “帮我把这篇文章做成一个课程视频,先检查一下现有素材。”
  • “Continue the video for the product launch idea up to the next approval point.”
  • “Set up my cloned voice profile first, then start today's video production.”

Requirements

  • Node.js, to run the bundled production scripts

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. 建立 video-brief.json
  2. 检查声音依赖
  3. 建立 content-decision.json
  4. 建立 video-content-plan.json
  5. 冻结内容和口播
  6. 生成本次声音
  7. 建立正式时序与 video-unit.json
  8. 预览、渲染与检查
  9. 记录真实状态

What it can do on your machine

Read from SKILL.md and the folder at commit 1e6d869. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

AI Video Production Assistant loads about 924 tokens when it runs, and up to ~4.6k if it reads all its reference files. Until then it costs about 45 tokens; SKILL.md has 177 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~924
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from wanghui2323/ai-video-maker at commit 1e6d869, republished under its MIT licence (© wanghui2323). 177 words, ~924 tokens.

Download SKILL.mdSave it as .claude/skills/make-ai-video/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.
name
make-ai-video
description
协助把想法、文章、提纲、已有口播、资料包或音频制作成有来源、可审核、可恢复的 AI 视频,并可选接入本地声音克隆。适用于用户提出“帮我把这个思路做成视频”“文章转视频”“整理口播和分镜”“用我的声音生成”“制作课程视频或短视频”“检查字幕、渲染或发布状态”等任务;也适用于需要区分首次声纹建档、本次配音、视频渲染和人工发布门禁的场景。

AI 视频制作助手

从用户已有的想法或素材开始,先判断输入和当前阶段,再协助完成内容、声音、画面与审核。不要把文章当作默认入口,也不要把首次声音建档、本次正式配音和整片发布合并成一次生成。

启动即执行

触发 Skill 后,立即进入工作流并执行当前可安全完成的步骤,不要只复述流程、列菜单或让用户自己拼命令。

  1. 先运行 node scripts/doctor.mjs --json 检查基础能力。定位或创建本次生产包,检查已有文件、设备、运行时和当前状态;没有生产包时运行 node scripts/create-package.mjs --dir <目录> --input-mode <类型> --summary <摘要>,不要临时拼接一套不可复用目录。
  2. 把用户现有输入写入 video-brief.json,明确缺失项,并继续执行到下一个真实人工门禁。
  3. 可逆的本地检查、建目录、生成合同对象和校验应主动执行。安装依赖、下载大模型或需要额外系统权限时,说明体积、目录和用途后发起所需批准;获准后继续,不要退回成教程。
  4. 只有授权/权利不明、核心事实待核验、口播确认、声音所有者选声、整片审核和发布授权等人工门禁才暂停。
  5. 每次暂停或交付都报告:当前阶段 / 已完成动作 / 产物路径 / 校验结果 / 需要用户做的一个决定 / 批准后下一步。

用户说“帮我做成视频”代表持续推进到下一个人工门禁;用户明确要求本人克隆声音且没有可用 Profile 时,按下文时钟 A 先执行本地声音建档,然后自动回到时钟 B。不要把“给一句开始话术”当作完成。

读取所需合同

  1. 先读 references/input-routing.md,把用户输入整理为 VideoBrief。
  2. 设计内容时读 references/content-contract.md。
  3. 选择画幅和分镜时读 references/visual-routing.md。
  4. 使用声音、渲染或报告状态前读 references/production-gates.md。
  5. 只有用户要求克隆或复用克隆声音时,才读 references/voice-cloning.md。
  6. 进入画面与渲染阶段时读 references/rendering-adapter.md。如果当前项目没有渲染适配器,继续交付内容、声音、字幕和 video-unit.json,但不得声称已经可以生成正式 MP4。
  7. assets/example-package/ 只用于理解完整对象和测试,不作为新项目直接改写;需要首次本地声纹建档时,再复制 assets/voice-clone-starter/。

使用两条时钟

时钟 A:低频声音能力建立

仅在用户明确选择克隆声音且没有可用 VoiceProfile 时执行:

text
授权与私有范围
→ 参考录音和准确逐字稿
→ 三个校准候选
→ 机器 QA
→ 声音所有者选择
→ production-pilot VoiceProfile

这条链通常只在首次建档、参考或模型漂移、授权范围变化时重跑。它先于完整视频生产,但不生成本次正式口播。

时钟 B:每条视频自己的生产

每个视频项目都执行:

text
VideoBrief
→ ContentDecision
→ VideoContentPlan
→ 口播人工确认
→ 本次 VoiceRun 或其他配音
→ 正式时序
→ 视觉预检
→ 渲染与技术检查
→ 整片人工审核
→ 发布候选

如果已有可用声音档案,在项目开始时做一次 preflight,口播确认后再生成本次三个候选。不要在口播未冻结时提前生成正式音频。

执行工作流

1. 建立 video-brief.json

识别 idea、article、outline、script、source-pack 或 audio 输入。记录受众、期望变化、来源、权利、证据成熟度、时长/画幅偏好和声音意图。

用户只有想法时,协助展开方向,但把假设与待核验事实写入 verificationNeeds;不要编造个人经历、数据或案例。

2. 检查声音依赖
  • none、human 或普通 synthetic:按相应合同继续。
  • cloned 且已有 Profile:立即 preflight;漂移则阻断。
  • cloned 且没有 Profile:读取 references/voice-cloning.md,先做设备与本地模型预检;缺少模型时按该参考执行下载与锁定,随后走时钟 A,再自动进入完整生产。
3. 建立 content-decision.json

根据输入成熟度生成 2–6 个候选,每个候选只回答一个问题,包含核心判断、观众动作、来源锚点、待核验项、时长建议和画幅建议。一次只推荐一个;保留其余候选。

4. 建立 video-content-plan.json

按内容密度选择时长:

  • quick:35–50 秒,一个纠偏或动作;
  • standard:60–85 秒,一个机制、比较或诊断;
  • deep:90–120 秒,一个带证据的案例;
  • course-master:130–180 秒,只用于一个需要 3–5 个相互依赖章节的问题。

每段绑定来源、时间预算、口播职责、情绪和唯一视觉任务。先建立具体冲突,再引入术语;结尾给出动作或边界。

5. 冻结内容和口播

向用户展示主问题、核心判断、事实边界、时长、画幅和完整口播。content_plan_reviewed 与 narration_reviewed 未通过时,不生成本次克隆配音,也不做完整渲染。

6. 生成本次声音

克隆声音时先重新 preflight,再按认知章节整段生成三个候选。机器淘汰坏音频;声音所有者明确选择其中一个。本次选择只批准该 VoiceRun,不批准整片或发布。

7. 建立正式时序与 video-unit.json

把已审内容和选定声音翻译成连续 beat。每个 beat 只使用一个 visualJob:conflict、mechanism、comparison、evidence、action 或 conclusion。

估算字幕只能用于草案。使用转写或人工对齐后,才能把 timing_ready 标为通过。

8. 预览、渲染与检查

先检查当前项目是否存在可用的 Remotion、剪辑工程或其他渲染适配器。存在时,先检查开场、核心解释、证据/动作和结尾代表帧,再渲染完整视频;检查尺寸、帧率、编码、音频、字幕范围、缺失资产和隐私残留。不存在时,明确报告 rendering_adapter_required,保留已完成产物,不得把视觉计划写成已经生成成片。

9. 记录真实状态

分别记录 local_package、review_candidate、release_candidate,以及各平台 draft、previewed、scheduled、published。自动校验、文件存在、本人选声和整片发布是不同事实。

校验生产包

在 Skill 目录运行:

bash
node scripts/validate-package.mjs --dir <生产包>

加 --json 获取机器可读结果。修复错误后再推进状态;警告必须在交付说明中解释。

停止条件

遇到以下情况时停止并报告精确阻塞点:

  • 用户意图、输入类型或权利边界不明确;
  • 待核验事实会改变核心结论;
  • 请求时长会删除结论边界;
  • 声音授权、Profile、模型、参考或哈希漂移;
  • 口播未审却准备生成本次克隆音频;
  • 估算字幕被当成正式时序;
  • 渲染命令结束但目标文件不存在;
  • 自动检查被用来推断人工批准。

保留已完成对象,从失败阶段恢复,不要重做整条链。

© wanghui2323, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 31 other files (scripts, references, assets) in make-ai-video of wanghui2323/ai-video-maker.

  • SKILL.md
  • agents/openai.yaml
  • assets/example-package/captions.json
  • assets/example-package/content-decision.json
  • assets/example-package/narration-and-rhythm.md
  • assets/example-package/source-notes.md
  • assets/example-package/video-brief.json
  • assets/example-package/video-content-plan.json
  • assets/example-package/video-unit.json
  • assets/example-package/workflow-state.json
  • assets/voice-clone-starter/gitignore.snippet
  • assets/voice-clone-starter/narration.txt
  • assets/voice-clone-starter/reference-transcript.txt
  • assets/voice-clone-starter/voice-consent.json
  • assets/voice-clone-starter/voice-provider.json
  • references/content-contract.md
  • … and 16 more

Open the folder on GitHubat commit 1e6d869

Compare with similar skills

AI Video Production Assistant next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

AI Video Production Assistant compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
AI Video Production Assistant this skillwanghui2323/ai-video-maker101—~924Automated safety check: PassMIT
Ergo Remotion Videoitwanger/toBeBetterJavaer18k—~1.1kAutomated safety check: PassNone
Media Genclacky-ai/openclacky1.2k—~7.3kAutomated safety check: PassMIT
Stage EditOrkas-AI/Orkas-VideoStudio499—~2.4kAutomated safety check: PassMIT
Pneuma Clipcraftpandazki/pneuma-skills161—~7.5kAutomated safety check: NotesMIT
Clean Audiohassancs91/claude-youtube-editor328—~1.8kAutomated safety check: PassMIT

Similar skills

  • Ergo Remotion Video

    itwanger/toBeBetterJavaer

    把口播稿做成二哥风格的 Remotion 视频,包括整理视频用稿、火山 TTS 配音、音画对齐、逐章动画预览和导出带配音的 MP4。用户说“做视频”“口播稿转视频”“Remotion”“继续做下一章”“出片”“渲染”“改读音”“配音读错了”,或给出 docs/src/ai/video/ 下的稿子要做成视频时使用。共享工具、配置和素材在…

    18k GitHub stars~1.1k tokensUpdated today
    Media & CreativeAuto-check passed
  • Media Gen

    clacky-ai/openclacky

    Generate or edit images, videos, or audio in the current task.

    1.2k GitHub stars~7.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    499 GitHub stars~2.4k tokensUpdated 18 days ago
    Media & CreativeAuto-check passed
  • Pneuma Clipcraft

    pandazki/pneuma-skills

    AI-orchestrated video production on @pneuma-craft. An agent skill from pandazki/pneuma-skills.

    161 GitHub stars~7.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Clean Audio

    hassancs91/claude-youtube-editor

    Voice/audio cleanup step of the AI Video Editor pipeline — diagnose a video's background noise, pick the right denoise method, and produce a cleaned master (voice isolated, levels preserved, video…

    328 GitHub stars~1.8k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video Director

    Jamailar/Beav

    Canonical entrypoint for every AI chat request that asks to make, generate, plan, or edit a video, including promotional films, ads, short videos, product videos, reference-image videos…

    1.8k GitHub stars~9.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed

Questions about AI Video Production Assistant

What does AI Video Production Assistant do?

Turns an idea, article, outline or audio file into a sourced, reviewable AI video, tracking whether narration uses a human, synthetic or cloned voice. Runs a diagnostic script on activation to check what already exists in the current production package, then writes whatever the user has provided into a video brief and keeps moving forward until it reaches the next real approval point that needs a human decision, rather than just listing steps or handing commands back to the user. Reversible local actions such as checking files, creating folders, and validating the brief proceed on their own; installing dependencies or downloading a large voice model requires explaining size and purpose first and getting approval.

When should I use AI Video Production Assistant?

AI Video Production Assistant fits situations like: turning a written idea or article into a planned video with narration and visuals; checking current production status before resuming an in-progress video; setting up a cloned voice profile before using it in a video; reviewing a locked content plan and narration before rendering starts.

How do I install AI Video Production Assistant in Claude Code?

Run `npx skills add wanghui2323/ai-video-maker --skill make-ai-video -a claude-code`. Or copy the skill folder (make-ai-video in wanghui2323/ai-video-maker) into .claude/skills/make-ai-video in your project. Claude Code loads it when a task matches its description.

How do I install AI Video Production Assistant in Codex?

Run `npx skills add wanghui2323/ai-video-maker --skill make-ai-video -a codex`. Or copy the skill folder (make-ai-video in wanghui2323/ai-video-maker) into .agents/skills/make-ai-video in your project. Codex loads it when a task matches its description.

Can I use AI Video Production Assistant in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wanghui2323/ai-video-maker --skill make-ai-video -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/make-ai-video, .gemini/skills/make-ai-video, .github/skills/make-ai-video and .opencode/skills/make-ai-video in your project.

What does AI Video Production Assistant need to run?

Going by SKILL.md and its folder, AI Video Production Assistant needs the command-line tools its instructions call (node). Our summary lists: Node.js, to run the bundled production scripts.

Does AI Video Production Assistant access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is AI Video Production Assistant safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does AI Video Production Assistant use?

AI Video Production Assistant is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does AI Video Production Assistant use?

About 924 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.7k tokens, read only when the agent opens those files.

What are the alternatives to AI Video Production Assistant?

Skills that share tags, products or a category with AI Video Production Assistant: Ergo Remotion Video (itwanger/toBeBetterJavaer, 18k stars), Media Gen (clacky-ai/openclacky, 1.2k stars), Stage Edit (Orkas-AI/Orkas-VideoStudio, 499 stars) and Pneuma Clipcraft (pandazki/pneuma-skills, 161 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains AI Video Production Assistant?

wanghui2323 (a GitHub user) maintains it in wanghui2323/ai-video-maker, which has 101 GitHub stars. The repository was last updated on August 16, 2026.

Source: wanghui2323/ai-video-maker on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.