Agent skill

Minimax Voice Director

by bozhouDev in bozhouDev/video-skills-toolkit

用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。

MITAuto-check: notesMedia & Creative

Install Minimax Voice Director

skills CLI
$ npx skills add bozhouDev/video-skills-toolkit --skill minimax-voice-director -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bozhouDev/video-skills-toolkit minimax-voice-director --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bozhouDev/video-skills-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/minimax-voice-director .claude/skills/minimax-voice-director && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
minimax-voice-director
GitHub stars
150
Token cost
~667 tokens
SKILL.md length
178 words
Files
25 (incl. scripts, references, assets)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。

  • Works in 3 steps: 先导演,再审批 → 生成、选 Take,再批准音频 → 只从定稿音频生成字幕
  • Tasks that involve Text to speech and voice
  • SKILL.md covers 路由边界, 三阶段主流程, 引擎能力边界 and 停止条件
  • Runs Python scripts from its folder

What it does

Minimax Voice Director is an agent skill from bozhouDev/video-skills-toolkit. 用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。

Its SKILL.md is about 670 tokens, which your agent loads only when the skill is triggered. The skill folder holds 28 other files, including scripts, reference files and assets (for example `agents/openai.yaml`, `assets/minimax_tts.py` and `assets/voice-direction.template.yaml`).

It sits in Media & Creative, covering Text to speech and voice and Transcription. It works with MiniMax. The repository describes itself as: Video skills toolkit for Remotion talking-head, sketch story, and audio-to-subtitles workflows. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice
  • Tasks that involve Transcription

Example prompts

  • “/minimax-voice-director”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. 先导演,再审批
  2. 生成、选 Take,再批准音频
  3. 只从定稿音频生成字幕

What it can do on your machine

Read from SKILL.md and the folder at commit 4766a16. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Minimax Voice Director loads about 667 tokens when it runs, and up to ~5k if it reads all its reference files. Until then it costs about 50 tokens; SKILL.md has 178 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~667
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:34
    flow.md` 和 `references/runtime.md`,加载项目 `.env.r2`,不输出任何密钥或 Voice ID。

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bozhouDev/video-skills-toolkit at commit 4766a16, republished under its MIT licence (© bozhouDev). 178 words, ~667 tokens.

Download SKILL.mdSave it as .claude/skills/minimax-voice-director/SKILL.md (or your agent's skills folder). This skill also uses 24 other files; get the full folder from GitHub.
name
minimax-voice-director
description
用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。
metadata.tags
minimax, tts, voice-director, voiceover, subtitles, cloud-api

MiniMax 声音导演

把口播稿变成可审批、可复现、可局部返工的声音表演。不得把本 skill 缩减为“文本进、MP3 出”的单次 TTS 调用。

路由边界

  • 用户明确选择 MiniMax 配音,或当前项目已经用 MiniMax:进入三阶段主流程。
  • 用户只说 TTS、旁白、自己的声音,但没有选 MiniMax:不自动替换引擎,先根据上下文路由到合适的本地或云端工作流。
  • 用户要 MiniMax Voice ID、声音克隆、声音设计或旧函数兼容:读 references/runtime.md,该分支不自动触发配音主流程。
  • 音乐、BGM、歌曲或 cover:改用同级 music(ElevenLabs Music API)。
  • 不因其他 TTS 失败而静默转用 MiniMax;先说明云端费用、声音差异和数据上传。

三阶段主流程

1. 先导演,再审批
  1. 定义原稿和字幕文本真值,先处理产品名、人名、数字、缩写和多音字。
  2. 读 references/directing.md 设计全片基线、言语动作、潜台词、情绪弧线、节奏曲线、焦点词、语调和必要的声音事件。
  3. 读 references/schema.md,从 assets/voice-direction.template.yaml 创建 work/tts/voice-direction.yaml。平台私有标签不得出现在导演中间层。
  4. 运行 scripts/validate_direction.py 和 scripts/compile_direction.py,产生审查视图、lint、MiniMax manifest 和干净字幕稿。
  5. 用独立审查回合检查导演逻辑,重点展示改写、强焦点、强转折、长停顿、声音事件、局部参数和多 Take 段落;通过后运行 scripts/mark_reviewed.py 固化审查 hash。
  6. 停在这里等待用户批准。 用户没有明确批准时,不得运行 approve_direction.py,不得调用 MiniMax。
  7. 用户批准后运行 scripts/approve_direction.py,再重新编译。render_allowed 必须为 true。
2. 生成、选 Take,再批准音频
  1. 读 references/workflow.md 和 references/runtime.md,加载项目 .env.r2,不输出任何密钥或 Voice ID。
  2. 运行 scripts/render_segments.py。它必须同时验证人工批准、lint 和内容 hash;任一失效就拒绝云端生成。
  3. 正式生成优先使用长连续块:开头约 40 秒一块,后续约 3 分钟一块;时长只是软目标,必须在句号或完整意群结尾切分,绝不从句子中间硬截。每块默认 1 个 Take,局部问题只重生局部。
  4. 听审候选并写入 take-selection.yaml,然后运行 scripts/select_takes.py 和 scripts/finalize_voice.py。不得把原始候选或未重建停顿的拼接音频冒充最终成果。
  5. 在最终速度下验收开头 15 秒、一个中段、所有主要转折和收尾;检查发音、节奏、焦点、情绪连续性、音色漂移和声音标签执行。
  6. 再次停下等待用户批准最终音频。 获得批准后运行 scripts/approve_audio.py 和 scripts/publish_voice.py。
3. 只从定稿音频生成字幕
  1. 只对 audio_approved 且 hash 匹配的最终 WAV 生成时间轴;M4A/MP3 是发布衍生物,不作为字幕时间基准。任何 Take、速度、裁剪或拼接变化都使旧时间轴失效。
  2. 使用 audio-to-subtitles 生成 ASR 时间轴。该步骤需要 R2/MediaKit 凭证并会上传最终音频;使用项目已授权凭证,不打印密钥。
  3. ASR 只提供时间轴。用 work/tts/subtitle-source.txt 回填最终显示文本,不显示呼吸、停顿和发音标记。
  4. 读 references/output-layout.md 交付 SRT/VTT/JSON 和 raw ASR 证据,通过 scripts/mark_subtitled.py 把交付文件的 hash 绑定到已批准音频。数字人、口型和正式剪辑都必须以这份定稿音频为唯一时间基准。

引擎能力边界

导演稿使用 sound tags、局部参数或要求重音/语调控制时,读 references/capabilities.md。每个导演意图必须被标识为:

  • 引擎原生执行;
  • 通过口语改写、分段、整段参数和 Take 挑选近似执行;
  • 或引擎不支持。

不得把词级重音、局部音高 contour 或未验证的 Speech 2.8 emotion 字段宣称为精确控制。

停止条件

只有同时满足下列条件才能报告完成:

  • 导演稿、审查视图、lint、render manifest 和字幕干净稿已归档。
  • 最终 Take 和最终速度音频已经人工批准。
  • 正式 WAV/M4A/MP3 已发布,旧成果已可恢复备份。
  • 字幕来自定稿音频,显示文本已回填为干净原稿。
  • 开头、中段、主要转折和收尾已听审,没有未记录的局部返工。

© bozhouDev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 24 other files (scripts, references, assets) in skills/minimax-voice-director of bozhouDev/video-skills-toolkit.

  • SKILL.md
  • agents/openai.yaml
  • assets/minimax_tts.py
  • assets/voice-direction.template.yaml
  • references/capabilities.md
  • references/directing.md
  • references/output-layout.md
  • references/runtime.md
  • references/schema.md
  • references/workflow.md
  • scripts/approve_audio.py
  • scripts/approve_direction.py
  • scripts/compile_direction.py
  • scripts/direction_contract.py
  • scripts/finalize_segments.py
  • scripts/finalize_voice.py
  • scripts/mark_reviewed.py
  • … and 8 more

Open the folder on GitHubat commit 4766a16

Compare with similar skills

Minimax Voice Director next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Minimax Voice Director compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Minimax Voice Director this skillbozhouDev/video-skills-toolkit150—~667Automated safety check: NotesMIT
Vox ExplainerCK42BB/vox-explainer-skill109—~2.6kAutomated safety check: PassMIT
Videohub Story Editorcacity/VideoHub168—~2.3kAutomated safety check: NotesMIT
KrillinAI CLI Operatorkrillinai/OpenCreator13k—~869Automated safety check: PassApache-2.0
Tts SkillPluviobyte/rnskill1.6k—~1.3kAutomated safety check: NotesCustom licence
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Vox Explainer

    CK42BB/vox-explainer-skill

    End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.

    109 GitHub stars~2.6k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Videohub Story Editor

    cacity/VideoHub

    把长视频或已有字幕转成有完整叙事的几分钟短片。先基于原文字幕和画面证据理解、选段与重排,再对最终时间轴重新翻译和可选润色;既可输出保留原声的双语字幕版,也可把原声降到 30% 并用 MiniMax 或豆包 TTS 生成影视解说、短剧混剪、播客串讲或知识解读版。已有项目可进入本地五轨时间线继续调整切点、旁白、原声窗口、字幕、音量和转场,并按修订版本渲染。用于“把长视频讲成短故事”“按字幕自动剪辑”…

    168 GitHub stars~2.3k tokensUpdated 9 days ago
    Media & CreativeAuto-check: notes
  • KrillinAI CLI Operator

    krillinai/OpenCreator

    Routes agents to the right KrillinAI command for subtitles, dubbing, video rendering, covers and speech, and explains how to read its JSON and manifest output.

    13k GitHub stars~869 tokensUpdated today
    Media & CreativeAuto-check passed
  • Tts Skill

    Pluviobyte/rnskill

    Generate cloned narration for the content workspace. An agent skill from Pluviobyte/rnskill.

    1.6k GitHub stars~1.3k tokensUpdated 20 days ago
    Media & CreativeAuto-check: notes
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Elevenlabs Transcribe

    qdhenry/Claude-Command-Suite

    Transcribes audio/video files using ElevenLabs Scribe v2 API.

    1.3k GitHub stars~1.5k tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes

More from bozhouDev/video-skills-toolkit

All 12 skills in this repo
  • Media To Transcript

    bozhouDev/video-skills-toolkit

    Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0.

    150 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check: notes
  • Audio To Subtitles

    bozhouDev/video-skills-toolkit

    Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API.

    150 GitHub stars~1.9k tokensUpdated 2 mo ago
    Auto-check: notes
  • Talking Head Hyperframes

    bozhouDev/video-skills-toolkit

    为 HyperFrames 口播或旁白项目创建、修复并验证固定舞台,锁定数字人 PIP 的区域、裁切、人物安全区和不透明背景,归档输入,生成 manifest 与 template handoff,并在就绪后按“字幕驱动的全镜头静态审核→动效”门禁路由到 hyperframes-scene-animator。适用于“新建 HyperFrames…

    150 GitHub stars~927 tokensUpdated 2 mo ago
    Auto-check passed
  • Music

    bozhouDev/video-skills-toolkit

    Generate music using ElevenLabs Music API. An agent skill from bozhouDev/video-skills-toolkit.

    150 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Viral Video Benchmark

    bozhouDev/video-skills-toolkit

    判断、扫描、拆解并归档抖音视频、小红书图文或小红书视频。实时读取用户同平台粉丝数并划分主对标池/跨级灵感池,用已登录浏览器读取目标作品和作者主页公开指标,再用确定性代码判定普通、小爆、爆款、现象级并扫描作者近 20 条候选;只对用户选中的爆款和现象级先构建可追溯证据包,再调用子 Agent…

    150 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Douyin Cover

    bozhouDev/video-skills-toolkit

    生成抖音、视频号、小红书等短视频封面图、视频标题图和合集封面,也能诊断和改版已有封面。用户说做封面、生成封面、抖音封面、视频封面、标题图、合集封面、3:4、4:3、1:1、短视频首图、动态封面首帧、给这期视频做图、这封面为什么没人点、帮我改封面、封面点击率怎么提升、诊断封面时都应使用。小白学AI系列封面除外:遇到“小白学AI封面/小白学AI第N集封面”时优先使用…

    150 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Minimax Voice Director

What does Minimax Voice Director do?

用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。. Minimax Voice Director is an agent skill from bozhouDev/video-skills-toolkit.

When should I use Minimax Voice Director?

Minimax Voice Director fits situations like: tasks that involve Text to speech and voice; tasks that involve Transcription.

How do I install Minimax Voice Director in Claude Code?

Run `npx skills add bozhouDev/video-skills-toolkit --skill minimax-voice-director -a claude-code`. Or copy the skill folder (skills/minimax-voice-director in bozhouDev/video-skills-toolkit) into .claude/skills/minimax-voice-director in your project. Claude Code loads it when a task matches its description.

How do I install Minimax Voice Director in Codex?

Run `npx skills add bozhouDev/video-skills-toolkit --skill minimax-voice-director -a codex`. Or copy the skill folder (skills/minimax-voice-director in bozhouDev/video-skills-toolkit) into .agents/skills/minimax-voice-director in your project. Codex loads it when a task matches its description.

Can I use Minimax Voice Director in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bozhouDev/video-skills-toolkit --skill minimax-voice-director -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/minimax-voice-director, .gemini/skills/minimax-voice-director, .github/skills/minimax-voice-director and .opencode/skills/minimax-voice-director in your project.

What does Minimax Voice Director need to run?

Going by SKILL.md and its folder, Minimax Voice Director needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Minimax Voice Director access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Minimax Voice Director safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Minimax Voice Director use?

Minimax Voice Director is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Minimax Voice Director use?

About 667 tokens (SKILL.md is roughly 2.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.3k tokens, read only when the agent opens those files.

What are the alternatives to Minimax Voice Director?

Skills that share tags, products or a category with Minimax Voice Director: Vox Explainer (CK42BB/vox-explainer-skill, 109 stars), Videohub Story Editor (cacity/VideoHub, 168 stars), KrillinAI CLI Operator (krillinai/OpenCreator, 13k stars) and Tts Skill (Pluviobyte/rnskill, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Minimax Voice Director?

bozhouDev (a GitHub user) maintains it in bozhouDev/video-skills-toolkit, which has 150 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on July 27, 2026.

Source: bozhouDev/video-skills-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.