Agent skill

Audio Ducking

by sonilo-ai in sonilo-ai/skills

Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.

MITAuto-check: notesMedia & Creative

Install Audio Ducking

skills CLI
$ npx skills add sonilo-ai/skills --skill audio-ducking -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sonilo-ai/skills audio-ducking --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/audio-ducking .claude/skills/audio-ducking && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio-ducking
GitHub stars
115
Token cost
~1.8k tokens
SKILL.md length
722 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.

  • Works in 3 steps: Sonilo MCP tools visible in this session… → No usable Sonilo MCP tools, but sonilo… → Neither — stop and run the setup-api-key…
  • Mixing a separately-generated
  • SKILL.md covers Transport: MCP or CLI, Quick Start, Tool and Parameters, plus 4 more sections
  • Calls pip, npm and curl; reaches api.sonilo.com; needs SONILO_API_KEY

What it does

Audio Ducking is an agent skill from sonilo-ai/skills. Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps. Use when mixing a separately-generated or existing music track under narration, dialogue, or a video's own voice track, without manual volume automation.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires Sonilo through either transport — the MCP server connected, or the sonilo CLI installed and signed in — plus credentials: a sonilo login sign-in, the…

It sits in Media & Creative, covering Text to speech and voice and MCP servers. It works with Model Context Protocol. The repository describes itself as: Agent skills for Sonilo's licensed music, sound-effects, dubbing, and audio-ducking API. The licence is MIT.

When your agent uses it

  • Mixing a separately-generated
  • Existing music track under narration
  • A videos own voice track
  • Without manual volume automation

Example prompts

  • “/audio-ducking”

Requirements

  • Python 3
  • Node.js
  • A credential in SONILO_API_KEY
  • Compatibility (from SKILL.md): Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.
  • Pre-approved tools (allowed-tools): Bash, Read, Write, mcp__sonilo__*

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Sonilo MCP tools visible in this session (audio_ducking and friends) — use them. This is the preferred path: it needs no shell, and it is…
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe…
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads…

What it can do on your machine

Read from SKILL.md and the folder at commit 1ce1bd8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • mcp__sonilo__*

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • npm
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.sonilo.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SONILO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.

    From compatibility in the SKILL.md frontmatter.

Context cost

Audio Ducking loads about 1.8k tokens when it runs. Until then it costs about 77 tokens; SKILL.md has 722 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, mcp__sonilo__*

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sonilo-ai/skills at commit 1ce1bd8, republished under its MIT licence (© sonilo-ai). 722 words, ~1,752 tokens.

Download SKILL.mdSave it as .claude/skills/audio-ducking/SKILL.md (or your agent's skills folder).
name
audio-ducking
description
Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps. Use when mixing a separately-generated or existing music track under narration, dialogue, or a video's own voice track, without manual volume automation.
allowed-tools
Bash, Read, Write, mcp__sonilo__*
compatibility
Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.
license
MIT

Sonilo Audio Ducking

Automatically duck a music bed under a voice track: Sonilo lowers the music wherever the voice is speaking and lifts it back in the gaps, then returns the mixed result. The voice input may be a video — its audio track is used as the voice, and the ducked mix is muxed back into a new video.

Setup: See the setup-api-key skill.

⚠️ Cost: makes an API call that may incur charges. Only call when explicitly requested.

Transport: MCP or CLI

Pick one at the start of the session and stay on it. Do not mix the two inside a single job, and do not announce the choice.

  1. Sonilo MCP tools visible in this session (audio_ducking and friends) — use them. This is the preferred path: it needs no shell, and it is the only one that survives a very long generation. If a call fails to authenticate — rather than failing on its inputs — this transport is not usable in this session: go to 2 instead of retrying it.
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe with sonilo account, not sonilo whoami: whoami exits 0 even when signed out, so it cannot tell the two states apart.
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads, polling and retries that a bare request does not.

Quick Start

audio_ducking(
    voice_path="~/Desktop/interview.mp4",
    music_path="~/Desktop/background-track.wav"
)
Python (pip install "sonilo>=0.13")
python
from sonilo import Sonilo

client = Sonilo()  # reads SONILO_API_KEY

result = client.audio_ducking.generate(
    voice="interview.mp4",  # audio or video; also voice_url=
    music="background-track.wav",  # audio only; also music_url=
)
result.save("ducked.mp4" if result.output_type == "video" else "ducked.wav")
JavaScript / TypeScript (npm install sonilo@>=0.14)
ts
import { SoniloClient, download } from "sonilo";
import { writeFile } from "node:fs/promises";

const client = new SoniloClient(); // reads SONILO_API_KEY

const result = await client.audioDucking.generate({
  voice: "./interview.mp4", // audio or video; also voiceUrl
  musicUrl: "https://example.com/background-track.wav", // audio only; also music
});
await writeFile(
  result.output_type === "video" ? "ducked.mp4" : "ducked.wav",
  await download(result.output_url!),
);
CLI (npm install -g sonilo-cli or pip install sonilo-cli)
bash
sonilo audio-ducking --voice interview.mp4 --music-url https://example.com/background-track.wav

Always async under the hood — the CLI submits and polls for you. Exactly one of --voice/--voice-url and one of --music/--music-url. The default output name follows what comes back (output.wav, or output.mp4 when the voice input was a video); --output overrides it. A local --music file must have an audio extension — the CLI rejects a video there up front, for the same reason the MCP tool does.

cURL (raw REST API, no MCP host)
bash
curl -X POST "https://api.sonilo.com/v1/audio-ducking" \
  -H "Authorization: Bearer $SONILO_API_KEY" \
  -F "voice_file=@interview.mp4" \
  -F "music_file=@background-track.wav"
# -> {"task_id": "..."}  poll GET /v1/tasks/{task_id}

A local file uses the voice_file/music_file multipart fields; a remote source uses voice_url/music_url form fields instead (mix and match freely between the two inputs).

Tool

ToolDescription
audio_ducking(voice_path? | voice_url?, music_path? | music_url?, output_directory?)Mix music under voice, ducking automatically wherever the voice speaks.

Parameters

ParameterTypeNotes
voice_pathstringAbsolute path, or relative to SONILO_MCP_BASE_PATH. Audio or video: .wav/.mp3/.m4a/.aac/.ogg/.flac or .mp4/.mov/.avi/.wmv/.webm/.mkv.
voice_urlstringHTTPS URL to the voice audio/video. Exactly one of voice_path/voice_url.
music_pathstringAbsolute path, or relative to the base path. Audio only — a video here is not treated specially and will be mishandled.
music_urlstringHTTPS URL to the music audio. Exactly one of music_path/music_url.
output_directorystringDefaults to SONILO_MCP_BASE_PATH.

Each input is capped at 360 seconds (6 minutes) and by the account's upload-size limit (typically 300 MB).

Show full SKILL.md (249 more words)Show less

Workflow Tips

  • This tool takes two already-existing tracks — it does not generate music or SFX itself. If you need to generate the music bed first, use the text-to-music or video-to-music skill (text_to_music/video_to_music), then feed the result in here as music_path.
  • The voice input can be a video. If the user hands you a talking-head clip or an interview and a separate music file, pass the video straight through as voice_path — Sonilo extracts its audio track, ducks the music under it, and re-muxes the ducked mix back into a new video automatically.
  • Prefer video-to-sound or video_to_music(ducking=true) when the music itself is also being generated for that same video — those tools duck internally as part of generation, so you don't need a separate ducking call. Reach for audio_ducking specifically when the music track is fixed/external and you just need the mix.

Recovering a Timed-Out Call

This tool submits an async task on the backend. If the call times out, the error carries a task_id — the job keeps running (already charged). Call get_sfx_task(task_id) later (get_generation_task(task_id) on the hosted server); see task-recovery.

Output Files

A single file: a .wav if the voice input was audio, or a .mp4 (ducked mix re-muxed in) if the voice input was a video. Named after the voice input (e.g. interview.mp4 → interview-ducked.mp4), falling back to ducked-<first 8 chars of the task id>.

Error Handling

Common errors: 401 invalid key, 402 insufficient balance / trial exhausted, 413 file too large, 422 invalid parameters, 429 rate limit. See the account skill.

© sonilo-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in audio-ducking of sonilo-ai/skills.

Open the folder on GitHubat commit 1ce1bd8

Compare with similar skills

Audio Ducking next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio Ducking compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio Ducking this skillsonilo-ai/skills115—~1.8kAutomated safety check: NotesMIT
Neurolink Guidejuspay/neurolink144—~1.4kAutomated safety check: PassMIT
Resolve Audiosamuelgursky/davinci-resolve-mcp3.4k—~1.3kAutomated safety check: PassMIT
GitHub Commentingjuspay/neurolink144—~698Automated safety check: PassMIT
Voiceguaardvark/guaardvark257—~798Automated safety check: PassMIT
Generate Narration AudioArcReel/ArcReel5.4k—~524Automated safety check: PassAGPL-3.0

Similar skills

  • Neurolink Guide

    juspay/neurolink

    Guide for using the NeuroLink SDK and CLI. An agent skill from juspay/neurolink.

    144 GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Resolve Audio

    samuelgursky/davinci-resolve-mcp

    Audio and Fairlight work in the DaVinci Resolve MCP. An agent skill from samuelgursky/davinci-resolve-mcp.

    3.4k GitHub stars~1.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • GitHub Commenting

    juspay/neurolink

    How to post clean, rich, deduplicated GitHub PR review comments — suggestion blocks, multi-line anchors, markers, formatting rules.

    144 GitHub stars~698 tokensUpdated today
    DevelopmentAuto-check passed
  • Voice

    guaardvark/guaardvark

    Narration and text-to-speech on the user's machine through Guaardvark's Audio Foundry (Chatterbox, Kokoro, Piper) and consent-gated voice cloning from a reference clip.

    257 GitHub stars~798 tokensUpdated today
    Media & CreativeAuto-check passed
  • 为旁白/解说剧本逐分镜生成旁白配音(TTS)。当 TTS 项目首轮自动剪辑需要补齐缺失配音、用户要求生成或重新生成某个分镜或某集旁白配音,或批量配音中断需要补齐时使用。

    5.4k GitHub stars~524 tokensUpdated today
    Media & CreativeAuto-check passed
  • Scenario Audio

    scenario-labs/skills

    A skill your agent uses when generating or handling audio on Scenario via MCP.

    931 GitHub stars~3k tokensUpdated today
    Media & CreativeAuto-check passed

More from sonilo-ai/skills

All 13 skills in this repo
  • Text To Sfx

    sonilo-ai/skills

    Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match.

    115 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • Video To Music

    sonilo-ai/skills

    Score a video with original music using Sonilo — the model watches the cut and matches pacing, motion, and emotion, returning either the audio or a new video with the score muxed in.

    115 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Account

    sonilo-ai/skills

    Check the Sonilo account's available services, rate limits, free-trial allowance, and usage/billing history.

    115 GitHub stars~1.4k tokensUpdated today
    Auto-check: notes
  • Audio Playback

    sonilo-ai/skills

    Play a local audio file through the system's default speakers using Sonilo's MCP server.

    115 GitHub stars~813 tokensUpdated today
    Auto-check: notes
  • Auto Dubbing

    sonilo-ai/skills

    Dub a video into one or more other languages using Sonilo, translating and re-voicing the speech into a new video per language.

    115 GitHub stars~4.5k tokensUpdated today
    Auto-check: notes
  • Proofread

    sonilo-ai/skills

    Transcribe a video with Sonilo and translate the transcript into editable .srt files — one per target language, plus the detected source language — so the wording can be read and corrected before…

    115 GitHub stars~4.4k tokensUpdated today
    Auto-check: notes

Questions about Audio Ducking

What does Audio Ducking do?

Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps. Audio Ducking is an agent skill from sonilo-ai/skills. Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.

When should I use Audio Ducking?

Audio Ducking fits situations like: mixing a separately-generated; existing music track under narration; A videos own voice track; without manual volume automation.

How do I install Audio Ducking in Claude Code?

Run `npx skills add sonilo-ai/skills --skill audio-ducking -a claude-code`. Or copy the skill folder (audio-ducking in sonilo-ai/skills) into .claude/skills/audio-ducking in your project. Claude Code loads it when a task matches its description.

How do I install Audio Ducking in Codex?

Run `npx skills add sonilo-ai/skills --skill audio-ducking -a codex`. Or copy the skill folder (audio-ducking in sonilo-ai/skills) into .agents/skills/audio-ducking in your project. Codex loads it when a task matches its description.

Can I use Audio Ducking in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sonilo-ai/skills --skill audio-ducking -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-ducking, .gemini/skills/audio-ducking, .github/skills/audio-ducking and .opencode/skills/audio-ducking in your project.

What does Audio Ducking need to run?

Going by SKILL.md and its folder, Audio Ducking needs the command-line tools its instructions call (pip, npm and curl) and credentials named SONILO_API_KEY. Our summary lists: Python 3; Node.js; A credential in SONILO_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, mcp__sonilo__*. Compatibility (from SKILL.md): Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill..

Does Audio Ducking access the network?

SKILL.md names 1 domain. In commands or code: api.sonilo.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Audio Ducking safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Audio Ducking use?

Audio Ducking is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio Ducking use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio Ducking?

Skills that share tags, products or a category with Audio Ducking: Neurolink Guide (juspay/neurolink, 144 stars), Resolve Audio (samuelgursky/davinci-resolve-mcp, 3.4k stars), GitHub Commenting (juspay/neurolink, 144 stars) and Voice (guaardvark/guaardvark, 257 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio Ducking?

sonilo-ai (a GitHub organization) maintains it in sonilo-ai/skills, which has 115 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 9, 2026.

Source: sonilo-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.