Vox Explainer
CK42BB/vox-explainer-skill
End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.
Enables voice synthesis, voice cloning, voice design, and audio post-processing using MiniMax Voice API and FFmpeg.
$ npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills mm-voice-maker --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mm-voice-maker .claude/skills/mm-voice-maker && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "mm-voice-maker" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/mm-voice-maker into .claude/skills/mm-voice-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm-voice-maker", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/mm-voice-makerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills mm-voice-maker --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/mm-voice-maker .agents/skills/mm-voice-maker && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "mm-voice-maker" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/mm-voice-maker into .agents/skills/mm-voice-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm-voice-maker", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills mm-voice-maker --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/mm-voice-maker .cursor/skills/mm-voice-maker && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "mm-voice-maker" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/mm-voice-maker into .cursor/skills/mm-voice-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm-voice-maker", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/mm-voice-maker--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills mm-voice-maker --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/mm-voice-maker .gemini/skills/mm-voice-maker && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "mm-voice-maker" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/mm-voice-maker into .gemini/skills/mm-voice-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm-voice-maker", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills mm-voice-makerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/mm-voice-maker .github/skills/mm-voice-maker && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "mm-voice-maker" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/mm-voice-maker into .github/skills/mm-voice-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm-voice-maker", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills mm-voice-maker --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/mm-voice-maker .opencode/skills/mm-voice-maker && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "mm-voice-maker" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/mm-voice-maker into .opencode/skills/mm-voice-maker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "mm-voice-maker", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
mm-voice-makerEnables voice synthesis, voice cloning, voice design, and audio post-processing using MiniMax Voice API and FFmpeg.
Mm Voice Maker is an agent skill from LeoYeAI/openclaw-master-skills. Enables voice synthesis, voice cloning, voice design, and audio post-processing using MiniMax Voice API and FFmpeg. Use when converting text to speech, creating custom voices, or processing/merging audio.
Its SKILL.md is about 6.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including scripts (for example `_meta.json`, `check_environment.py` and `mmvoice.py`).
It sits in Media & Creative, covering Text to speech and voice and Video production. It works with MiniMax and FFmpeg. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 4 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
pythonpipbrewaptFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
MINIMAX_VOICE_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Mm Voice Maker loads about 6.6k tokens when it runs. Until then it costs about 55 tokens; SKILL.md has 1,725 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
stall: `brew install ffmpeg` (macOS) or `sudo apt install ffmpeg` (Ubuntu)Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 1,725 words, ~6,566 tokens.
.claude/skills/mm-voice-maker/SKILL.md (or your agent's skills folder). This skill also uses 23 other files; get the full folder from GitHub.Professional text-to-speech skill with emotion detection, voice cloning, and audio processing capabilities powered by MiniMax Voice API and FFmpeg.
| Area | Features |
|---|---|
| TTS | Sync (HTTP/WebSocket), async (long text), streaming |
| Segment-based | Multi-voice, multi-emotion synthesis from segments.json, auto merge |
| Voice | Cloning (10s–5min), design (text prompt), management |
| Audio | Format conversion, merge, normalize, trim, remove silence (FFmpeg) |
mmVoice_Maker/
├── SKILL.md # This overview
├── mmvoice.py # CLI tool (recommended for Agents)
├── check_environment.py # Environment verification
├── requirements.txt
├── scripts/ # Entry: scripts/__init__.py
│ ├── utils.py # Config, data classes
│ ├── sync_tts.py # HTTP/WebSocket TTS
│ ├── async_tts.py # Long text TTS
│ ├── segment_tts.py # Segment-based TTS (multi-voice, multi-emotion)
│ ├── voice_clone.py # Voice cloning
│ ├── voice_design.py # Voice design
│ ├── voice_management.py # List/delete voices
│ └── audio_processing.py # FFmpeg audio tools
└── reference/ # Load as needed
├── cli-guide.md # CLI usage guide
├── getting-started.md # Setup and quick test
├── tts-guide.md # Sync/async TTS workflows
├── voice-guide.md # Clone/design/manage
├── audio-guide.md # Audio processing
├── script-examples.md # Runnable code snippets
├── troubleshooting.md # Common issues
├── api_documentation.md # Complete API reference
└── voice_catalog.md # Voice selection guide6-step workflow: [step1]. Verify environment
[step2-preparation]⚠️NOTE: Before processing the text, you must read voice-catalog.md for voice selection.
[step2]. Process text into script → <cwd>/audio/segments.json. Note: [Step2.4] is really important, you must check it twice before sending the script to the user.
[step2.5]. ⚠️ Generate preview for user confirmation (highly recommended for multi-voice content)
[step3]. Present plan to user for confirmation
[step4]. Validate segments.json
[step5]. Generate and merge audio → intermediate files in <cwd>/audio/tmp/, final output in <cwd>/audio/output.mp3
[step6]. ⚠️ CRITICAL: User confirms audio quality FIRST → THEN cleanup temp files (only after user is satisfied)
<cwd>is Claude's current working directory (not the skill directory). Audio files are saved relative to where Claude is running commands.
python check_environment.pyChecks:
If API key is not set, ask user for keys and set it:
export MINIMAX_VOICE_API_KEY="your-api-key-here"⚠️ MOST IMPORTANT PRINCIPLE: Gender Matching First
Before selecting voices, you MUST always match gender first. This is non-negotiable.
Golden Rule:
If a character is male → use male voice If a character is female → use female voice If a character is neutral/other → choose appropriate neutral voice
Why this matters:
Examples:
| Character | Wrong Voice | Correct Voice |
|---|---|---|
| 唐三藏 (male monk) | female-yujie ❌ | Chinese (Mandarin)_Gentleman ✅ |
| 林黛玉 (female) | male-qn-badao ❌ | female-shaonv ✅ |
| 曹操 (male warlord) | female-chengshu ❌ | Chinese (Mandarin)_Unrestrained_Young_Man ✅ |
Decision guide: Evaluate based on:
emotion empty)Use case scenarios:
| Scenario | Description | Segments | Voice Selection |
|---|---|---|---|
| Single Voice | User needs one voice for the entire content. Segment only by length (≤1,000,000 chars per segment). | Split by length only | One voice_id for all segments |
| Multi-Voice | Multiple characters/speakers, each with different voice. Segment by speaker/role changes. | Split by logical unit (speaker, dialogue, etc.) | Different voice_id per role |
| Podcast/Interview | Host and guest speakers with distinct voices. | Split by speaker | Voice per host/guest |
| Audiobook/Fiction | Narrator and character voices. | Split by narration vs. dialogue | Voice per narrator/character |
| Documentary | Mostly narration with occasional quotes. | Keep as one segment | Single narrator voice |
| Report/Announcement | Formal content with consistent tone. | Keep as one segment | Professional voice |
Processing Workflow (4 sub-steps):
Step 2.1: Text Segmentation and Role Analysis First, segment your text into logical units and identify the role/character for each segment.
Key principle (Important!): Split by logical unit, NOT simply by sentence
When to split (Important!):
When NOT to split (Important!):
Decision depends on use case:
| Use case | Example | Split strategy |
|---|---|---|
| Single Voice | Long article, news piece, announcement | Split by length (≤1,000,000 chars), same voice for all |
| Podcast/Interview | "Host: Welcome to the show. Guest: Thank you for having me." | Split by speaker |
| Documentary narration | "The scientist explained, 'The results are promising.'" | Keep as one segment (narrator voice) |
| Audiobook/Fiction | "'Who's there?' she whispered." | Split: "'Who's there?'" should be in character voice, while "she whispered." should be in narrator's voice |
| Report | "According to the report, the economy is growing." | Keep as one segment |
Example1: Single Voice (speech-2.8) For single-voice content (e.g., news, announcements, articles), segment only by length while maintaining the same voice:
[
{"text": "First part of the article (under 1,000,000 chars)...", "role": "narrator", "voice_id": "female-shaonv", "emotion": ""},
{"text": "Second part of the article (under 1,000,000 chars)...", "role": "narrator", "voice_id": "female-shaonv", "emotion": ""},
{"text": "Third part of the article (under 1,000,000 chars)...", "role": "narrator", "voice_id": "female-shaonv", "emotion": ""}
]Example2: Audiobook with characters (speech-2.8) In audiobooks (multi-voice fiction), split when narration and dialogue mix in the same sentence:
[
{"text": "The detective entered the room.", "role": "narrator", "voice_id": "", "emotion": ""},
{"text": "\"Who's there?\"", "role": "female_character", "voice_id": "", "emotion": ""},
{"text": "she whispered.", "role": "narrator", "voice_id": "", "emotion": ""},
{"text": "\"It's me,\"", "role": "male_character", "voice_id": "", "emotion": ""},
{"text": "he replied calmly.", "role": "narrator", "voice_id": "", "emotion": ""}
]Example3: Documentary/podcast narration (speech-2.8) Quoted speech in narration stays in narrator's voice (no need to split):
[
{
"text": "The scientist explained, \"The results show significant improvement in all test groups.\"",
"role": "narrator",
"voice_id": "",
"emotion": ""
},
{
"text": "According to the latest report, the economy has grown by 3% this quarter.",
"role": "narrator",
"voice_id": "",
"emotion": ""
}
]
**Note:** In the preliminary `segments.json`:
- Fill in the `text` field with segment content
- Fill in the `role` field to identify the character (narrator, male_character, female_character, host, guest, etc.)
- Leave `voice_id` empty (to be filled in Step 2.2)
- Leave `emotion` empty for speech-2.8 models
**Step 2.2: Voice Selection**
After segmenting and labeling roles, analyze all detected characters in your text. Consult [voice_catalog.md](reference/voice_catalog.md) **Section 1 "How to Choose a Voice"** to match voices to characters.
**⚠️ CRITICAL: Follow the two-step selection process below**
**Path A — Professional domains (Story/Narration, News/Announcements, Documentary):**
If the content belongs to one of these three professional domains, prioritize selecting from the recommended voices in **voice_catalog.md Section 2.1** (filter by scenario + gender). These voices are specifically optimized for their professional use cases.
**Path B — All other scenarios:**
Select from **voice_catalog.md Section 2.2**, following this strict priority hierarchy:
1. **First: Match Gender** (non-negotiable) — Male characters MUST use male voices, female characters MUST use female voices
2. **Second: Match Language** — The voice MUST match the content language (Chinese content → Chinese voice, Korean content → Korean voice, English content → English voice, etc.). Never assign a voice from the wrong language.
3. **Third: Match Age** — Determine the age group (Children / Youth / Adult / Elderly / Professional) and select from the corresponding subsection in Section 2.2
4. **Fourth: Match Personality & Role** — Choose the best fit based on personality traits, tone, and character role
**Voice Selection Decision Tree:**Is this a professional domain (Story/News/Documentary)? ├── YES → Select from voice_catalog Section 2.1 (filter by scenario + gender) └── NO → Select from voice_catalog Section 2.2: Step 1: Match Gender ├── Male character → Male voices only └── Female character → Female voices only Step 2: Match Age Group └── Children / Youth / Adult / Elderly / Professional Step 3: Match Language └── Filter to voices matching the content language Step 4: Match Personality & Role └── Choose best fit by tone, personality, character role
**Step 2.3: Emotions Segmentation** *(For non-2.8 series models only)*
For models other than speech-2.8 series, analyze emotions in your segments:
- For **long segments**, split further based on **emotional transitions**
- Add appropriate **emotion tags** to each segment
- Refer to Section 3 in [text-processing.md](reference/text-processing.md) for emotion tags and examples
- Skip this step for speech-2.8 models (emotion is auto-matched)
**Emotion Tags:**
- For speech-2.6 series (speech-2.6-hd and speech-2.6-turbo): happy, sad, angry, fearful, disgusted, surprised, calm, fluent, whisper
- For older models: happy, sad, angry, fearful, disgusted, surprised, calm (7 emotions)
**Step 2.4: Check and Post-processing**
Finally, review and optimize your script:
- Verify segment length limits (async TTS ≤1,000,000 characters)
- Clean up conversational text (remove speaker names if needed)
- Ensure consistency in voice and emotion tags
- **Critical check for multi-voice content**: For audiobooks, multi-voice fiction, or content where dialogue is presented from a first-person perspective, verify that narration and dialogue mixed in the same sentence are properly split.
**When splitting IS needed (first-person dialogue in fiction/audiobooks):**
Example: `"John asked, 'Where are you going?'"` should be split into:
- Segment 1: `"John asked, "` - uses narrator voice (describes who is speaking)
- Segment 2: `"Where are you going?"` - uses the character's voice (actual dialogue in first-person)
This ensures proper voice differentiation: descriptive narration uses the narrator's voice, while the character's spoken words use the character's designated voice.
**When splitting is NOT needed (third-person quotes in podcast/documentary/news):**
In podcasts, documentaries, or news reports, quoted speech is typically presented in third-person narrative style - the speaker's words are being reported, not performed. Keep these as one segment with the narrator's voice and remove the speaker's name at the beginning:
- `"Welcome to our show." → narrator voice, remove the speaker's name (like "The host said:") at the beginning
- `"According to experts, 'This technology represents a significant breakthrough.'" → keep as one segment (narrator voice)
- `"Scientists noted, 'The experimental results exceeded our expectations.'" → keep as one segment (narrator voice)
- **If the split is missing**: Go back to Step 2.1 and ensure dialogue portions are separated from narration with appropriate role labels.
**Create segments.json:**
After completing all 4 sub-steps, save the final `segments.json` to `<cwd>/audio/segments.json`.
### Step 2.5: Generate Preview for User Confirmation (Highly Recommended)
**For multi-voice content (audiobooks, dramas, etc.), always generate a preview first.**
This saves time and prevents waste when voice selections need adjustment.
**How to generate a preview:**
1. Create a smaller segments file with 10-20 representative segments (include all characters)
2. Generate the preview audio
3. Ask user to listen and confirm voice choices
**Preview segments.json example:**
```json
[
{"text": "Narration opening...", "role": "narrator", "voice_id": "...", "emotion": ""},
{"text": "Male character speaks...", "role": "male_character", "voice_id": "...", "emotion": ""},
{"text": "Female character speaks...", "role": "female_character", "voice_id": "...", "emotion": ""},
{"text": "More dialogue...", "role": "...", "voice_id": "...", "emotion": ""}
]Preview command:
python mmvoice.py generate segments_preview.json -o preview.mp3When user confirms preview:
Before proceeding to validation and generation, present the segmentation plan to the user and wait for confirmation:
Present to the user:
Example confirmation message:
I've analyzed the text and created a segmentation plan:
**Roles and Voices:**
- Narrator: male-qn-jingying (deep, authoritative, suitable for storytelling)
- Protagonist: female-shaonv (bright, energetic, youthful)
- Antagonist: male-qn-qingse (cool, menacing)
**Model:** speech-2.8-hd (recommended - automatic emotion matching)
**Language:** Chinese
**Segments:** 8 segments total
Please review and confirm:
1. ⚠️ **Gender Verification**: Do the voice genders match the character genders?
- [Narrator: Male ✓] [Protagonist: Female ✓] [Antagonist: Male ✓]
2. ⚠️ **Language Verification**: Do the voice languages match the content language?
- [All voices: Chinese ✓]
3. Are the voice assignments appropriate for each character (age, personality)?
4. Should any segments be combined or split differently?
5. Any other changes you'd like to make?
**After generation:**
- I'll generate a preview first for you to review
- Only after you confirm the audio quality will I clean up temporary files
- If not satisfied, I'll re-generate and we iterate until you're happy
Reply "confirm" to proceed, or let me know what to adjust.Wait for user response:
segments.json and present the plan again for confirmationBefore generating audio, validate the segments file:
# Default: speech-2.8-hd (auto emotion matching)
python mmvoice.py validate <cwd>/audio/segments.json
# Specify model for context-specific validation
python mmvoice.py validate <cwd>/audio/segments.json --model speech-2.6-hd
# Validate voice_ids against available voices (slower, requires API call)
python mmvoice.py validate <cwd>/audio/segments.json --validate-voices
# Combined options (recommended)
python mmvoice.py validate <cwd>/audio/segments.json --model speech-2.6-hd --validate-voices
# Use `--verbose` to see segment details
python mmvoice.py validate <cwd>/audio/segments.json --model speech-2.6-hd --validate-voices --verbose
Emotion Validation checks:
| Model | Emotion Validation |
|---|---|
| speech-2.8-hd/turbo | Emotion can be empty (auto emotion matching) |
| speech-2.6-hd/turbo | All 9 emotions supported |
| Older models | happy, sad, angry, fearful, disgusted, surprised, calm (7 emotions) |
Voice ID validation:
With --validate-voices:
Generate audio for all segments and merge into final output.
File placement (default behavior if user doesn't specify):
<cwd>/ # Claude's current working directory
└── audio/ # Created automatically
├── tmp/ # Intermediate segment files
│ ├── segment_0000.mp3
│ ├── segment_0001.mp3
│ └── ...
└── <custom_audio_name>.mp3 # Final merged audio, name can be customizedWhere <cwd> is Claude's current working directory (where commands are executed).
-o is not specified, output goes to <cwd>/audio/output.mp3<cwd>/audio/tmp/<cwd>/audio/tmp/Basic usage:
# Default: speech-2.8-hd, output to <cwd>/audio/output.mp3
python mmvoice.py generate <cwd>/audio/segments.json
# Specify output path
python mmvoice.py generate <cwd>/audio/segments.json -o <cwd>/audio/<custom_audio_name>.mp3
# Specify model if needed
python mmvoice.py generate <cwd>/audio/segments.json --model speech-2.6-hdSkip existing segments (for rate limit retries):
# Only generate segments that don't exist yet - skips already-generated files
python mmvoice.py generate <cwd>/audio/segments.json --skip-existingError handling:
--continue-on-error to generate remaining segments despite failures--skip-existing to skip already successfully generated segments (recommended for retries after rate limit)⚠️ CRITICAL: Never delete temp files until user confirms!
After generation completes, you MUST follow this exact sequence:
Step 6.1: Report generation result to user
✓ Audio saved to: <output_path>
Generated: X/Y segments
Intermediate files in: <cwd>/audio/tmp/Step 6.2: Ask user to confirm audio quality Ask the user to listen to the audio and confirm:
Step 6.3: Wait for user response
Step 6.4: Only after user confirms, offer cleanup
After confirming audio quality, temporary files can be deleted with:
rm -rf <cwd>/audio/tmp/NEVER execute rm -rf on temp files without explicit user confirmation!
If user is NOT satisfied:
Use the following when the task involves voice creation, single-voice TTS (sync/async), or audio processing instead of the main segment-based workflow. Each subsection gives CLI commands, script paths, and the reference doc to open for details.
mmvoice.py):python mmvoice.py clone AUDIO_FILE --voice-id VOICE_ID # Clone from 10s–5min audio
python mmvoice.py design "DESCRIPTION" --voice-id ID # Design from text
python mmvoice.py list-voices # List all voicesscripts/voice_clone.py (clone), scripts/voice_design.py (design), scripts/voice_management.py (list/manage).python mmvoice.py tts "TEXT" -o OUTPUT.mp3 [-v VOICE_ID] [--model MODEL]scripts/sync_tts.py (HTTP/WebSocket sync), scripts/async_tts.py (async task + poll).python mmvoice.py merge FILE1 [FILE2 ...] -o OUTPUT [--crossfade MS]
python mmvoice.py convert INPUT -o OUTPUT [--format FORMAT]scripts/audio_processing.py (merge, convert, normalize, trim).validate and generate as in Steps 4–5 above.scripts/segment_tts.py.Open these when you need concrete usage, parameters, or troubleshooting. Paths are relative to the skill root.
| Document | Content for the Agent |
|---|---|
| reference/cli-guide.md | All CLI commands (validate, generate, tts, clone, design, list-voices, merge, convert, check-env) with options and examples. Use for correct CLI invocation. |
| reference/getting-started.md | Environment setup (venv, pip install, FFmpeg), MINIMAX_VOICE_API_KEY, basic synthesis test. Use for first-time setup or “env not working”. |
| reference/tts-guide.md | Sync TTS (short text), async TTS (long text), streaming TTS, multi-segment production. Use for sync/async/streaming logic and parameters. |
| reference/voice-guide.md | Voice cloning (quick, high-quality with prompt audio, step-by-step), voice design, voice management. Use for custom voice creation flows. |
| reference/audio-guide.md | Format conversion, merging (including crossfade and fallback), normalization, trimming, optimization. Use for merge/convert/normalize behavior and options. |
| reference/script-examples.md | Copy-paste runnable examples for sync TTS, async TTS, segment-based TTS, audio processing, voice clone/design/management. Use for quick Python snippets. |
| reference/troubleshooting.md | Environment (API key, FFmpeg), API errors, segment-based TTS, audio, voice. Use when an error message or unexpected behavior appears. |
| reference/api_documentation.md | Full API reference: config, sync/async TTS, emotion parameter, segment-based TTS, voice clone/design/management, audio processing, common parameters, error handling. Use for exact function signatures and parameter details. |
| reference/voice_catalog.md | System voices list (male/female/beta), selection guide, voice parameters, custom voices, voice IDs. Use to choose or look up voice_id. |
MINIMAX_VOICE_API_KEY environment variable must be setbrew install ffmpeg (macOS) or sudo apt install ffmpeg (Ubuntu)<#x#> in text where x = pause duration in seconds (0.01–99.99)"Hello<#1.5#>world" creates 1.5s pause between wordspython check_environment.py to diagnose setup issues© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 23 other files (scripts) in skills/mm-voice-maker of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Mm Voice Maker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Mm Voice Maker this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~6.6k | Automated safety check: Notes | MIT | |
| Vox ExplainerCK42BB/vox-explainer-skill | 109 | — | ~2.6k | Automated safety check: Pass | MIT | |
| Videohub Story Editorcacity/VideoHub | 168 | — | ~2.3k | Automated safety check: Notes | MIT | |
| MiniMax Multimodal Toolkitmadebyaris/advance-minimax-m3-cursor-rules | 126 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Media Productionleon-ai/leon | 18k | — | ~999 | Automated safety check: Pass | MIT | |
| Vox DirectorAlisa0808/vox-director | 2.2k | — | ~5.6k | Automated safety check: Pass | MIT |
CK42BB/vox-explainer-skill
End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.
cacity/VideoHub
把长视频或已有字幕转成有完整叙事的几分钟短片。先基于原文字幕和画面证据理解、选段与重排,再对最终时间轴重新翻译和可选润色;既可输出保留原声的双语字幕版,也可把原声降到 30% 并用 MiniMax 或豆包 TTS 生成影视解说、短剧混剪、播客串讲或知识解读版。已有项目可进入本地五轨时间线继续调整切点、旁白、原声窗口、字幕、音量和转场,并按修订版本渲染。用于“把长视频讲成短故事”“按字幕自动剪辑”…
madebyaris/advance-minimax-m3-cursor-rules
Routes image, video, voice, music and media-processing requests to the smallest suitable MiniMax or local path, with FFmpeg-style processing around generated media.
leon-ai/leon
Generates images, audio and video through Leon's media tools, joins them with FFmpeg, checks the output and attaches playable files for the owner.
Alisa0808/vox-director
Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…
limin112/min-skill
Build a narrated explainer video from a concept — discussion → structure → HTML slide deck → narration script → TTS voice → subtitles → background music → Playwright screen recording → ffmpeg…
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Categories
Enables voice synthesis, voice cloning, voice design, and audio post-processing using MiniMax Voice API and FFmpeg. Mm Voice Maker is an agent skill from LeoYeAI/openclaw-master-skills. Enables voice synthesis, voice cloning, voice design, and audio post-processing using MiniMax Voice API and FFmpeg.
Mm Voice Maker fits situations like: converting text to speech; creating custom voices; processing/merging audio.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a claude-code`. Or copy the skill folder (skills/mm-voice-maker in LeoYeAI/openclaw-master-skills) into .claude/skills/mm-voice-maker in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a codex`. Or copy the skill folder (skills/mm-voice-maker in LeoYeAI/openclaw-master-skills) into .agents/skills/mm-voice-maker in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill mm-voice-maker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mm-voice-maker, .gemini/skills/mm-voice-maker, .github/skills/mm-voice-maker and .opencode/skills/mm-voice-maker in your project.
Going by SKILL.md and its folder, Mm Voice Maker needs Python for the scripts in its folder, the command-line tools its instructions call (python, pip, brew and apt) and credentials named MINIMAX_VOICE_API_KEY. Our summary lists: Python 3; A credential in MINIMAX_VOICE_API_KEY.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Mm Voice Maker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.6k tokens (SKILL.md is roughly 26k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Mm Voice Maker: Vox Explainer (CK42BB/vox-explainer-skill, 109 stars), Videohub Story Editor (cacity/VideoHub, 168 stars), MiniMax Multimodal Toolkit (madebyaris/advance-minimax-m3-cursor-rules, 126 stars) and Media Production (leon-ai/leon, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.