Agent skill

Audio Mix

by ZJU-REAL in ZJU-REAL/Easel

音频混合 / 混音:把旁白口播 + 背景音乐 + 音效混成一轨,BGM 自动循环补足并可闪避(旁白说话时自动压低 BGM 保证人声清晰)。当用户说 混音、音频混合、旁白加背景音乐、配音加BGM、人声和音乐混一起、加音效、音频叠加、BGM 压低、闪避、ducking、把配音和bgm合起来 时使用。基于 shared/scripts/audiomix.py。与 audio-editing…

Apache-2.0Auto-check passedMedia & Creative

Install Audio Mix

skills CLI
$ npx skills add ZJU-REAL/Easel --skill audio-mix -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ZJU-REAL/Easel audio-mix --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ZJU-REAL/Easel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/openclaw/audio-mix .claude/skills/audio-mix && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio-mix
GitHub stars
3.4k
Token cost
~483 tokens
SKILL.md length
83 words
Files
1
Skills in repo
114
Repo updated
First seen
Licence
Apache-2.0

At a glance

音频混合 / 混音:把旁白口播 + 背景音乐 + 音效混成一轨,BGM 自动循环补足并可闪避(旁白说话时自动压低 BGM 保证人声清晰)。当用户说 混音、音频混合、旁白加背景音乐、配音加BGM、人声和音乐混一起、加音效、音频叠加、BGM 压低、闪避、ducking、把配音和bgm合起来 时使用。基于 shared/scripts/audiomix.py。与 audio-editing…

  • Works in 5 steps: 有旁白时输出时长 = 旁白长度,BGM 自动循环/裁切对齐并在末尾淡出。 → 旁白 + BGM 默认开启闪避(人声优先);不需要时显式 --no-duck。 → sfx 与 --sfx-at 数量一致(或不给 --sfx-at 全部默认 0s)。 → …
  • Tasks that involve Video production
  • SKILL.md covers 输入, 输出(outputs/主题名/), 执行步骤 and 调参, plus 2 more sections
  • Calls python

What it does

Audio Mix is an agent skill from ZJU-REAL/Easel. 音频混合 / 混音:把旁白口播 + 背景音乐 + 音效混成一轨,BGM 自动循环补足并可闪避(旁白说话时自动压低 BGM 保证人声清晰)。当用户说 混音、音频混合、旁白加背景音乐、配音加BGM、人声和音乐混一起、加音效、音频叠加、BGM 压低、闪避、ducking、把配音和bgm合起来 时使用。基于 shared/scripts/audiomix.py。与 audio-editing concat 区别:concat 是前后顺序拼接,本 SKILL 是同时叠加混音;与 video-editing bgm 区别:那个给视频配乐,本 SKILL 输出纯音频。

Its SKILL.md is about 480 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Video production and Music and audio generation. The repository describes itself as: An open-source AI agent for social media — discover trends, create content, publish everywhere, and learn what works across Xiaohongshu, Douyin, Zhihu, Bilibili, and more.🎨一个开源的… The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Video production
  • Tasks that involve Music and audio generation

Example prompts

  • “/audio-mix”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. 有旁白时输出时长 = 旁白长度,BGM 自动循环/裁切对齐并在末尾淡出。
  2. 旁白 + BGM 默认开启闪避(人声优先);不需要时显式 --no-duck。
  3. sfx 与 --sfx-at 数量一致(或不给 --sfx-at 全部默认 0s)。
  4. 混音不做响度归一(保留相对音量);需统一响度先用 audio-editing normalize。
  5. 产物统一进 outputs/主题名/。

What it can do on your machine

Read from SKILL.md and the folder at commit 278f420. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio Mix loads about 483 tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 83 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~483

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ZJU-REAL/Easel at commit 278f420, republished under its Apache-2.0 licence (© ZJU-REAL). 83 words, ~483 tokens.

Download SKILL.mdSave it as .claude/skills/audio-mix/SKILL.md (or your agent's skills folder).
name
audio-mix
description
音频混合 / 混音:把旁白口播 + 背景音乐 + 音效混成一轨,BGM 自动循环补足并可闪避(旁白说话时自动压低 BGM 保证人声清晰)。当用户说 混音、音频混合、旁白加背景音乐、配音加BGM、人声和音乐混一起、加音效、音频叠加、BGM 压低、闪避、ducking、把配音和bgm合起来 时使用。基于 shared/scripts/audio_mix.py。与 audio-editing concat 区别:concat 是前后顺序拼接,本 SKILL 是同时叠加混音;与 video-editing bgm 区别:那个给视频配乐,本 SKILL 输出纯音频。
layer
produce

音频混合(旁白 + BGM + 音效)

把多条音频同时叠加混成一轨,核心能力是闪避(ducking)——旁白说话时自动压低 背景音乐,人声清晰、音乐不抢。全部走 skills/shared/scripts/audio_mix.py, 不要手拼 amix/sidechaincompress。

前后顺序拼接(一段接一段)见 audio-editing concat;给视频配乐见 video-editing bgm; 降噪见 audio-denoise。

输入

字段必填说明
旁白否口播/配音主轨(给了则输出时长跟它走,并触发闪避)
BGM否背景音乐(自动循环补足到旁白长度)
音效否一个或多个音效,可指定各自出现时间点

(三者至少给一个。最典型:"旁白 + BGM"。)

输出(outputs/主题名/)

  • 混音后的单轨音频(mp3/wav/m4a,按输出后缀)
  • 报告:轨数、时长、是否闪避

执行步骤

脚本路径(相对项目根):skills/shared/scripts/audio_mix.py(mix -h 看参数)。

bash
# 旁白 + BGM(默认自动闪避,BGM 循环补足到旁白长度)
python skills/shared/scripts/audio_mix.py mix \
  --voice narration.mp3 --bgm music.mp3 --bgm-volume 0.25 \
  -o outputs/主题名/final.mp3

# 关闭闪避(纯叠加)
python skills/shared/scripts/audio_mix.py mix --voice v.mp3 --bgm m.mp3 --no-duck -o out.mp3

# 旁白 + 定时音效(第 3.5s 一个叮,第 8s 一个 whoosh)
python skills/shared/scripts/audio_mix.py mix --voice v.mp3 \
  --sfx ding.wav --sfx-at 3.5 --sfx whoosh.wav --sfx-at 8 -o out.mp3

调参

  • 人声被音乐盖住:调低 --bgm-volume(默认 0.25)或确认闪避已开(默认开)。
  • 闪避太猛/音乐一顿一顿:--no-duck 后手动压低 --bgm-volume。
  • BGM 比旁白短:默认自动循环;不想循环用 --bgm-loop-off。
  • 音效太响/太轻:--sfx-volume(默认 0.9)。

规则

  1. 有旁白时输出时长 = 旁白长度,BGM 自动循环/裁切对齐并在末尾淡出。
  2. 旁白 + BGM 默认开启闪避(人声优先);不需要时显式 --no-duck。
  3. --sfx 与 --sfx-at 数量一致(或不给 --sfx-at 全部默认 0s)。
  4. 混音不做响度归一(保留相对音量);需统一响度先用 audio-editing normalize。
  5. 产物统一进 outputs/主题名/。

参考来源

闪避用 ffmpeg sidechaincompress(以人声为控制信号压缩 BGM),是播客/口播视频保证人声清晰的 标准做法;多轨叠加用 amix。把 sidechain 接线与循环对齐封装成确定性脚本。

© ZJU-REAL, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/openclaw/audio-mix of ZJU-REAL/Easel.

Open the folder on GitHubat commit 278f420

Compare with similar skills

Audio Mix next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio Mix compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio Mix this skillZJU-REAL/Easel3.4k—~483Automated safety check: PassApache-2.0
Music to Videoheygen-com/hyperframes60k3 repos~4.7kAutomated safety check: NotesApache-2.0
Paper Collage Explainer Generatortl2012tl/comfyUI-llama-TE2414 repos~5.2kAutomated safety check: PassNone
Short Form Editnateherkai/hyperframes-student-kit1.3k—~5.3kAutomated safety check: PassCustom licence
Acestepcalesthio/OpenMontage66k—~2.3kAutomated safety check: NotesAGPL-3.0
Tesseract Videomirage-hq/Tesseract130—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Music to Video

    heygen-com/hyperframes

    Turns a music track into a beat-synced HyperFrames video such as a lyric video, slideshow or kinetic promo, with any supplied images or clips cut onto the beat grid.

    60k GitHub starsUsed in 3 repos~4.7k tokens
    Media & CreativeAuto-check: notes
  • Paper Collage Explainer Generator

    tl2012tl/comfyUI-llama-TE

    For creators, educators, and social-video editors who need a tactile paper-collage language for narration, knowledge points, opinions, or abstract topics.

    241 GitHub starsUsed in 4 repos~5.2k tokens
    Media & CreativeAuto-check passed
  • Short Form Edit

    nateherkai/hyperframes-student-kit

    Turn talking-head footage into a finished reel, YouTube Short, or short advertisement with curiosity-led openings, earned payoffs, story-driven cuts, transcript-synced motion graphics, moving…

    1.3k GitHub stars~5.3k tokensUpdated 12 days ago
    Media & CreativeAuto-check passed
  • Acestep

    calesthio/OpenMontage

    AI music generation with ACE-Step 1.5 — background music, vocal tracks, covers, stem extraction for video production.

    66k GitHub stars~2.3k tokensUpdated 7 days ago
    Media & CreativeAuto-check: notes
  • Tesseract Video

    mirage-hq/Tesseract

    Edit existing footage into finished videos locally with Tesseract.

    130 GitHub stars~1.9k tokensUpdated 9 days ago
    Media & CreativeAuto-check passed
  • Make Video

    BintzGavin/helios

    Make videos drawn with code: motion graphics, launch and explainer videos, animated shorts, animated charts and data, logo reveals, social clips, music visualizers, UI demos and GIF loops.

    119 GitHub stars~2k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed

More from ZJU-REAL/Easel

All 114 skills in this repo
  • Gzh Design

    ZJU-REAL/Easel

    微信公众号文章排版引擎:把 Markdown / Word(.docx) / PDF / 纯文本转成可直接粘贴进公众号编辑器的 HTML,自动章节编号、关键词标记、引言卡、目录、代码块、图片/GIF、作者签名;主题从 references/theme-index.md…

    3.4k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • 微信公众号文章自动创作与发布工具。给定参考文章、文字或文档,自动搜索整理全网相关信息、生成图文并茂的公众号文章,并发布到微信公众号草稿箱。特别强调反 AI 检测写作。

    3.4k GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Card Design

    ZJU-REAL/Easel

    社媒卡片视觉设计系统:提供配色、中文字体层级、满画幅布局、品类骨架和死空白/密度质检,避免模板化 PPT 与廉价 AI 感。

    3.4k GitHub stars~657 tokensUpdated today
    Auto-check passed
  • Ecom Details Image

    ZJU-REAL/Easel

    生成电商商品视觉方案:主图概念、场景图、详情页视觉方向和 AI 生图 Prompt. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Infographic

    ZJU-REAL/Easel

    将数据或文字内容转化为可视化信息图,支持静态(AntV)和动画 GIF 两种模式。当用户需要制作信息图、数据可视化、流程图、对比图、动画图表、GIF 图表、思维导图、SWOT 分析图时调用。本地渲染信息图/GIF 动画;要单张静态图片 URL 用 chart-visualization,要 CSV/JSON→整页报告用 data-report

    3.4k GitHub stars~643 tokensUpdated today
    Auto-check passed
  • Novel Writer

    ZJU-REAL/Easel

    长篇小说/网文连载创作:从世界观、人设和三级大纲写到逐章正文,并用文件化状态维护伏笔、前情和跨章一致性. An agent skill from ZJU-REAL/Easel.

    3.4k GitHub stars~1k tokensUpdated today
    Auto-check passed

Questions about Audio Mix

What does Audio Mix do?

音频混合 / 混音:把旁白口播 + 背景音乐 + 音效混成一轨,BGM 自动循环补足并可闪避(旁白说话时自动压低 BGM 保证人声清晰)。当用户说 混音、音频混合、旁白加背景音乐、配音加BGM、人声和音乐混一起、加音效、音频叠加、BGM 压低、闪避、ducking、把配音和bgm合起来 时使用。基于 shared/scripts/audiomix.py。与 audio-editing…. Audio Mix is an agent skill from ZJU-REAL/Easel.

When should I use Audio Mix?

Audio Mix fits situations like: tasks that involve Video production; tasks that involve Music and audio generation.

How do I install Audio Mix in Claude Code?

Run `npx skills add ZJU-REAL/Easel --skill audio-mix -a claude-code`. Or copy the skill folder (skills/openclaw/audio-mix in ZJU-REAL/Easel) into .claude/skills/audio-mix in your project. Claude Code loads it when a task matches its description.

How do I install Audio Mix in Codex?

Run `npx skills add ZJU-REAL/Easel --skill audio-mix -a codex`. Or copy the skill folder (skills/openclaw/audio-mix in ZJU-REAL/Easel) into .agents/skills/audio-mix in your project. Codex loads it when a task matches its description.

Can I use Audio Mix in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ZJU-REAL/Easel --skill audio-mix -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-mix, .gemini/skills/audio-mix, .github/skills/audio-mix and .opencode/skills/audio-mix in your project.

What does Audio Mix need to run?

Going by SKILL.md and its folder, Audio Mix needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Audio Mix access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Audio Mix safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audio Mix use?

Audio Mix is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio Mix use?

About 483 tokens (SKILL.md is roughly 1.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio Mix?

Skills that share tags, products or a category with Audio Mix: Music to Video (heygen-com/hyperframes, 60k stars), Paper Collage Explainer Generator (tl2012tl/comfyUI-llama-TE, 241 stars), Short Form Edit (nateherkai/hyperframes-student-kit, 1.3k stars) and Acestep (calesthio/OpenMontage, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio Mix?

ZJU-REAL (a GitHub organization) maintains it in ZJU-REAL/Easel, which has 3,376 GitHub stars. The repository holds 114 skills in this directory. The repository was last updated on October 9, 2026.

Source: ZJU-REAL/Easel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.