Agent skill

Video Voiceover and Music Mixer

by SpaceZephyr in SpaceZephyr/creator-buddy

Adds sound to a finished video: generates narration from a script with TTS, finds royalty-free background music, and mixes everything together with ffmpeg.

No licenceAuto-check passedMedia & Creative

SKILL.md written in Chinese; this summary is our English description.

Install Video Voiceover and Music Mixer

skills CLI
$ npx skills add SpaceZephyr/creator-buddy --skill space-video-audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SpaceZephyr/creator-buddy space-video-audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SpaceZephyr/creator-buddy.git skills-src && mkdir -p .claude/skills && cp -r skills-src/video-Skills/space-video-audio .claude/skills/space-video-audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
space-video-audio
GitHub stars
1.6k
Token cost
~615 tokens
SKILL.md length
164 words
Files
4 (incl. scripts, references)
Skills in repo
31
Repo updated
First seen
Licence
None found

At a glance

Adds sound to a finished video: generates narration from a script with TTS, finds royalty-free background music, and mixes everything together with ffmpeg.

  • Works in 6 steps: 对齐长度:BGM 太短自动循环,太长裁切到视频时长 → 响度归一:loudnorm 把整体拉到 -14 LUFS 左右(平台友好) → 压音 ducking:有旁白时用 sidechaincompress,人声一响… → …
  • Adding a narrated voiceover to a video from a written script
  • SKILL.md covers 一、配音(TTS 旁白), 二、配乐(找 BGM 并下载), 三、混音(合成到视频) and 交接, plus 2 more sections
  • Runs Shell scripts from its folder; calls uvx, bash and curl

What it does

The skill covers three jobs that can be used separately or together: narration, background music and the final mix. For narration it picks a text-to-speech backend by tier. IndexTTS2 with a cloned voice is used when a local model is set up, edge-tts (run through uvx, with Chinese voices such as zh-CN-XiaoxiaoNeural and zh-CN-YunxiNeural) is the default, and macOS say is the offline fallback. Long scripts are generated in sections and joined, and subtitles are timed by re-transcribing the narration audio.

For music, the agent decides mood and tempo from the video type, then looks at royalty-free sources such as Pixabay Music, YouTube Audio Library, Free Music Archive, ccMixter and Incompetech, checking each licence and noting any required credit for the video description. Tracks are downloaded with curl or yt-dlp, copyrighted pop music is off limits, and the user decides when a licence is unclear.

Mixing runs through scripts/mix.sh, which takes the video, a music file, an optional voice file and an output name, plus options for music level, fade length and ducking. It loops or trims the music to the video length, normalizes loudness to about -14 LUFS, lowers the music under speech with sidechain compression, adds fades and copies the video stream without re-encoding.

When your agent uses it

  • Adding a narrated voiceover to a video from a written script
  • Choosing and downloading licence-safe background music that fits a video
  • Mixing music and narration into a video with ducking and loudness control

Example prompts

  • “Generate a voiceover from script.txt and add upbeat background music to demo.mp4.”
  • “Find royalty-free lo-fi music for my tutorial video and mix it quietly under the narration.”
  • “Mix bgm.mp3 and voice.mp3 into final.mp4 with ducking and a one second fade.”

Requirements

  • ffmpeg with the loudnorm and sidechaincompress filters
  • yt-dlp for downloading tracks
  • uvx to run edge-tts
  • An IndexTTS2 setup for the cloned-voice tier (optional)

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. 对齐长度:BGM 太短自动循环,太长裁切到视频时长
  2. 响度归一:loudnorm 把整体拉到 -14 LUFS 左右(平台友好)
  3. 压音 ducking:有旁白时用 sidechaincompress,人声一响 BGM 自动下沉,人声停 BGM 回来
  4. 淡入淡出:首尾 afade,别硬切
  5. BGM 音量:无旁白约 -12~-16dB,有旁白垫底约 -20~-26dB
  6. 封装:-c:v copy 不重编码画面,只替换/混合音轨

What it can do on your machine

Read from SKILL.md and the folder at commit edf46c5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • uvx
    • bash
    • curl
    • yt-dlp

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uvx and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Voiceover and Music Mixer loads about 615 tokens when it runs, and up to ~1.6k if it reads all its reference files. Until then it costs about 56 tokens; SKILL.md has 164 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~56
When it runs · the whole SKILL.md, loaded when a task matches
~615
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 164 words (~615 tokens).

name
space-video-audio

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (scripts, references) in video-Skills/space-video-audio of SpaceZephyr/creator-buddy.

  • SKILL.md
  • references/mixing-recipes.md
  • references/music-sources.md
  • scripts/mix.sh

Open the folder on GitHubat commit edf46c5

Compare with similar skills

Video Voiceover and Music Mixer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Voiceover and Music Mixer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Voiceover and Music Mixer this skillSpaceZephyr/creator-buddy1.6k—~615Automated safety check: PassNone
Qiaomu Cutjoeseesun/qiaomu-cut-skill372—~6.8kAutomated safety check: NotesMIT
ShowtimeFavioVazquez/showtime220—~3kAutomated safety check: PassMIT
Media Productionleon-ai/leon18k—~1kAutomated safety check: PassMIT
Release Videohuytieu/COG-second-brain1.3k—~1.7kAutomated safety check: PassMIT
Vox DirectorAlisa0808/vox-director2.2k—~5.6kAutomated safety check: PassMIT

Similar skills

  • Qiaomu Cut

    joeseesun/qiaomu-cut-skill

    把一句话需求转成可复现、可验收视频工程的乔木智能剪辑导演。Use when the user asks to create, plan, edit, remix, explain, narrate, subtitle, animate, composite, or render a video—including one-line requests such as “制作一个科普视频:介绍…

    372 GitHub stars~6.8k tokensUpdated 12 days ago
    Media & CreativeAuto-check: notes
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    220 GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Media Production

    leon-ai/leon

    Generates images, audio and video through Leon's media tools, joins them with FFmpeg, checks the output and attaches playable files for the owner.

    18k GitHub stars~1k tokensUpdated today
    Media & CreativeAuto-check passed
  • Release Video

    huytieu/COG-second-brain

    Turn a product release (the list of shipped items plus real screen recordings) into a motion recap video and one explained demo per feature, with sound effects tied to on-screen motion and a…

    1.3k GitHub stars~1.7k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Video Podcast Maker

    dtsola/xiaoyaosearch

    Turns a topic into a 4K horizontal video podcast through research, scripting, text-to-speech, Remotion rendering and background music, and can learn styles from references.

    1k GitHub stars~3.4k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed

More from SpaceZephyr/creator-buddy

All 31 skills in this repo
  • WeChat Hot Article Analysis

    SpaceZephyr/creator-buddy

    Fetches hot WeChat Official Account articles by sector or keywords and produces a data file and an HTML report with rankings, style patterns and writing references.

    1.6k GitHub stars~847 tokensUpdated 1 mo ago
    Auto-check passed
  • Xiaohongshu Multi-Page HTML Cards

    SpaceZephyr/creator-buddy

    Turns an article, SOP, or tutorial into a self-contained, multi-page 3:4 HTML deck of image cards meant to be screenshotted page by page.

    1.6k GitHub stars~928 tokensUpdated 1 mo ago
    Auto-check passed
  • WeChat Viral Article Finder

    SpaceZephyr/creator-buddy

    Finds recent viral WeChat official account articles for a niche keyword and shows up to ten picks as cards, for topic research and benchmarking.

    1.6k GitHub stars~848 tokensUpdated 1 mo ago
    Auto-check passed
  • Draws chart images such as flowcharts, architecture diagrams, ER diagrams and SWOT boards for WeChat official account articles, in one of six visual styles.

    1.6k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • WeChat Article Cover Maker

    SpaceZephyr/creator-buddy

    Creates a 2.35:1 header cover for a WeChat official account article with image generation, keeping the headline inside the area that survives sharing crops.

    1.6k GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Space Video Broll

    SpaceZephyr/creator-buddy

    动效导演 + B-roll 生成器。从一段文字/脚本/视频,一次性产出 30 秒以上的动效 B-roll 视频(HTML→确定性 MP4)。不做 PPT 式翻页,从「视觉隐喻 + 运动」出发;内置多种设计风格(暗色 SaaS 魔术、黑白打字机、暖色编辑、极简白、蓝图网格);产出一律不用 emoji,用排版/几何/SVG 线性图标。当用户说「做 B-roll」「做个动效视频」「motion…

    1.6k GitHub stars~716 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Video Voiceover and Music Mixer

What does Video Voiceover and Music Mixer do?

Adds sound to a finished video: generates narration from a script with TTS, finds royalty-free background music, and mixes everything together with ffmpeg. The skill covers three jobs that can be used separately or together: narration, background music and the final mix. For narration it picks a text-to-speech backend by tier.

When should I use Video Voiceover and Music Mixer?

Video Voiceover and Music Mixer fits situations like: adding a narrated voiceover to a video from a written script; choosing and downloading licence-safe background music that fits a video; mixing music and narration into a video with ducking and loudness control.

How do I install Video Voiceover and Music Mixer in Claude Code?

Run `npx skills add SpaceZephyr/creator-buddy --skill space-video-audio -a claude-code`. Or copy the skill folder (video-Skills/space-video-audio in SpaceZephyr/creator-buddy) into .claude/skills/space-video-audio in your project. Claude Code loads it when a task matches its description.

How do I install Video Voiceover and Music Mixer in Codex?

Run `npx skills add SpaceZephyr/creator-buddy --skill space-video-audio -a codex`. Or copy the skill folder (video-Skills/space-video-audio in SpaceZephyr/creator-buddy) into .agents/skills/space-video-audio in your project. Codex loads it when a task matches its description.

Can I use Video Voiceover and Music Mixer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SpaceZephyr/creator-buddy --skill space-video-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/space-video-audio, .gemini/skills/space-video-audio, .github/skills/space-video-audio and .opencode/skills/space-video-audio in your project.

What does Video Voiceover and Music Mixer need to run?

Going by SKILL.md and its folder, Video Voiceover and Music Mixer needs a shell for the scripts in its folder and the command-line tools its instructions call (uvx, bash, curl and yt-dlp). Our summary lists: ffmpeg with the loudnorm and sidechaincompress filters; yt-dlp for downloading tracks; uvx to run edge-tts; An IndexTTS2 setup for the cloned-voice tier (optional).

Does Video Voiceover and Music Mixer access the network?

SKILL.md contains no URLs. Its commands use uvx and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Video Voiceover and Music Mixer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Voiceover and Music Mixer use?

No licence was found for Video Voiceover and Music Mixer or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Video Voiceover and Music Mixer use?

About 615 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 964 tokens, read only when the agent opens those files.

What are the alternatives to Video Voiceover and Music Mixer?

Skills that share tags, products or a category with Video Voiceover and Music Mixer: Qiaomu Cut (joeseesun/qiaomu-cut-skill, 372 stars), Showtime (FavioVazquez/showtime, 220 stars), Media Production (leon-ai/leon, 18k stars) and Release Video (huytieu/COG-second-brain, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Voiceover and Music Mixer?

SpaceZephyr (a GitHub user) maintains it in SpaceZephyr/creator-buddy, which has 1,613 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on August 26, 2026.

Source: SpaceZephyr/creator-buddy on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.