Agent skill

Video To Sound

by sonilo-ai in sonilo-ai/skills

Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo.

MITAuto-check: notesMedia & Creative

Install Video To Sound

skills CLI
$ npx skills add sonilo-ai/skills --skill video-to-sound -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sonilo-ai/skills video-to-sound --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/video-to-sound .claude/skills/video-to-sound && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-to-sound
GitHub stars
115
Token cost
~2.6k tokens
SKILL.md length
1,096 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo.

  • Works in 3 steps: Sonilo MCP tools visible in this session… → No usable Sonilo MCP tools, but sonilo… → Neither — stop and run the setup-api-key…
  • Tasks that involve Music and audio generation
  • SKILL.md covers Transport: MCP or CLI, Quick Start, Tools and Parameters, plus 5 more sections
  • Calls pip, npm and curl; reaches api.sonilo.com; needs SONILO_API_KEY

What it does

Video To Sound is an agent skill from sonilo-ai/skills. Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo. Use instead of calling the music and sound-effects skills separately for the same video — the two layers are mixed and ducked against each other by the backend. Returns a mixed audio track, or a new video with it muxed in.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires Sonilo through either transport — the MCP server connected, or the sonilo CLI installed and signed in — plus credentials: a sonilo login sign-in, the…

It sits in Media & Creative, covering Music and audio generation. It works with Model Context Protocol. The repository describes itself as: Agent skills for Sonilo's licensed music, sound-effects, dubbing, and audio-ducking API. The licence is MIT.

When your agent uses it

  • Tasks that involve Music and audio generation

Example prompts

  • “/video-to-sound”

Requirements

  • Python 3
  • Node.js
  • A credential in SONILO_API_KEY
  • Compatibility (from SKILL.md): Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.
  • Pre-approved tools (allowed-tools): Bash, Read, Write, mcp__sonilo__*

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Sonilo MCP tools visible in this session (video_to_sound and friends) — use them. This is the preferred path: it needs no shell, and it is…
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe…
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads…

What it can do on your machine

Read from SKILL.md and the folder at commit 1ce1bd8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • mcp__sonilo__*

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip
    • npm
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.sonilo.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SONILO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.

    From compatibility in the SKILL.md frontmatter.

Context cost

Video To Sound loads about 2.6k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,096 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, mcp__sonilo__*

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sonilo-ai/skills at commit 1ce1bd8, republished under its MIT licence (© sonilo-ai). 1,096 words, ~2,628 tokens.

Download SKILL.mdSave it as .claude/skills/video-to-sound/SKILL.md (or your agent's skills folder).
name
video-to-sound
description
Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo. Use instead of calling the music and sound-effects skills separately for the same video — the two layers are mixed and ducked against each other by the backend. Returns a mixed audio track, or a new video with it muxed in.
allowed-tools
Bash, Read, Write, mcp__sonilo__*
compatibility
Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.
license
MIT

Sonilo Video-to-Sound (Music + SFX Combined)

Generate a music bed and sound effects for a video in one call, balanced against each other and mixed by the backend — one charge instead of two separate generations. Use this whenever a video needs a full soundtrack (score + SFX), not just one or the other.

Setup: See the setup-api-key skill.

⚠️ Cost: makes one API call that may incur charges (billed once, not twice, even though it produces both layers). Only call when explicitly requested.

Transport: MCP or CLI

Pick one at the start of the session and stay on it. Do not mix the two inside a single job, and do not announce the choice.

  1. Sonilo MCP tools visible in this session (video_to_sound and friends) — use them. This is the preferred path: it needs no shell, and it is the only one that survives a very long generation. If a call fails to authenticate — rather than failing on its inputs — this transport is not usable in this session: go to 2 instead of retrying it.
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe with sonilo account, not sonilo whoami: whoami exits 0 even when signed out, so it cannot tell the two states apart.
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads, polling and retries that a bare request does not.

Quick Start

video_to_sound(
    video_path="~/Desktop/trailer.mp4",
    music_prompt="Cinematic, building tension",
    sfx_prompt="Footsteps, wind, distant thunder"
)
video_to_video_sound(
    video_path="~/Desktop/trailer.mp4",
    music_prompt="Cinematic, building tension"
)
Python (pip install sonilo)
python
from sonilo import Sonilo

client = Sonilo()  # reads SONILO_API_KEY

mix = client.video_to_sound.generate(
    video="trailer.mp4",
    music_prompt="Cinematic, building tension",
    sfx_prompt="Footsteps, wind, distant thunder",
)
mix.save("soundtrack.wav")

video = client.video_to_video_sound.generate(video="trailer.mp4", music_prompt="Cinematic, building tension")
video.save("scored.mp4")
JavaScript / TypeScript (npm install sonilo)
ts
import { SoniloClient, download } from "sonilo";
import { writeFile } from "node:fs/promises";

const client = new SoniloClient(); // reads SONILO_API_KEY

const mix = await client.videoToSound.generate({
  video: "./trailer.mp4",
  musicPrompt: "Cinematic, building tension",
  sfxPrompt: "Footsteps, wind, distant thunder",
});
await writeFile("soundtrack.wav", await download(mix.output_url));

const video = await client.videoToVideoSound.generate({
  video: "./trailer.mp4",
  musicPrompt: "Cinematic, building tension",
});
await writeFile("scored.mp4", await download(video.output_url));
CLI (npm install -g sonilo-cli or pip install sonilo-cli)
bash
sonilo video-to-sound --video trailer.mp4 \
  --music-prompt "Cinematic, building tension" --sfx-prompt "Footsteps, wind, distant thunder" \
  --output soundtrack.wav

sonilo video-to-video-sound --video trailer.mp4 --music-prompt "Cinematic, building tension"

Unlike the music/sound-effects skills, both tools here have CLI commands. --stem music/--stem sfx (repeatable) additionally saves the individual layers next to the combined output.

cURL (raw REST API, no MCP host)
bash
curl -X POST "https://api.sonilo.com/v1/video-to-sound" \
  -H "Authorization: Bearer $SONILO_API_KEY" \
  -F "video=@trailer.mp4" \
  -F "music_prompt=Cinematic, building tension" \
  -F "sfx_prompt=Footsteps, wind, distant thunder"
# -> {"task_id": "..."}  poll GET /v1/tasks/{task_id}

Both endpoints are task-based (202 + poll), same as the sound-effects tools — the MCP tool waits for you.

Tools

ToolDescription
video_to_sound(video_path? | video_url?, music_prompt?, sfx_prompt?, segments?, preserve_speech?, ducking?, output_format?, variants_num?, output_directory?)Generate and mix music + SFX for a video, returns a single audio file.
video_to_video_sound(video_path? | video_url?, music_prompt?, sfx_prompt?, segments?, keep_original_sound?, preserve_speech?, ducking?, variants_num?, output_directory?)Same, but returns a new .mp4 with the mixed soundtrack muxed in. By default the source's own audio is dropped — see keep_original_sound.

Parameters

ParameterTypeDefaultNotes
video_pathstring—.mp4/.mov/.webm/.m4v/.gif (gif must be animated). Max 480s (8 min), subject to the account's upload-size cap.
video_urlstring—HTTPS/HTTP URL. Exactly one of video_path/video_url.
music_promptstring—Style hint for the music bed (max 2000 chars). Optional — omit to let Sonilo decide.
sfx_promptstring—Description of the SFX layered over the music (max 2000 chars). Optional.
segmentslist[dict]—Per-segment SFX descriptions — same schema and validation rules as in the video-to-sfx skill. Max 30 segments.
preserve_speechboolfalseKeep the source video's speech audible in the mix.
duckingboolfalseBrings the source video's own speech into the mix and dips the generated music under it. Off by default: with ducking and preserve_speech both unset, the result carries the generated music and effects alone and no music_processed stem exists. Pass true for any video with dialogue or narration that should stay audible.
keep_original_soundboolfalsevideo_to_video_sound only. Keeps the whole source track (dialogue, room tone, existing effects) with the generated mix over it, rather than replacing it. Add ducking=true to dip the mix under the voice instead of a flat blend. Supersedes preserve_speech.
output_formatstringwavvideo_to_sound only — video_to_video_sound always returns an .mp4. wav, m4a, or mp3 (320 kbps). Sets the combined track's container only; stems keep their own native formats.
variants_numint11–10 distinct mixes in one request, one file each. Cost scales linearly and any value above 1 is never free-trial covered — confirm the count with the user before calling.
output_directorystringSONILO_MCP_BASE_PATHAbsolute, or relative to the base path.
Show full SKILL.md (461 more words)Show less

Prompting

No prompt is required — the model reads the cut. A short structured brief adds your intent on top. Since this endpoint generates music and SFX in one balanced call, both crafts apply:

Workflow Tips

  • Use this instead of chaining video_to_music + video_to_sfx. The two layers are balanced against each other by the backend (so the SFX doesn't fight the score), and it's one charge, not two.
  • Both music_prompt and sfx_prompt are optional — you can leave both unset and let Sonilo interpret the whole scene, or set just one to steer that layer while leaving the other automatic.
  • ducking is off by default — turn it on for anything with a voice. Left off, the source speech is not in the mix at all: the output is generated music and effects only. That is the right default for a silent or music-only clip and the wrong one for a talking head, so check the source audio before calling (see the pre-flight reference) rather than after the user tells you the narration is gone.
  • For video_to_video_sound, the source audio is dropped unless you say otherwise. keep_original_sound=true keeps the whole original track under the generated mix; preserve_speech=true keeps only the isolated speech. If a user reports "my dialogue disappeared", this is the fix.
  • Only the combined mixed result is saved. The individual music/SFX/processed stems exist in the task body on the backend but are deliberately not downloaded — four files per call would bury the one the user actually wants. If stems are needed, call the REST API directly and inspect the task body.
  • Want the video back with the soundtrack baked in? Use video_to_video_sound instead of video_to_sound.
  • Don't know what it should sound like? Run video-analysis first: one call returns, by default, both a music section plan with ready-to-use generation prompts and a sound-design brief (sfx_segments + sfx_prompt) read off the footage, which beats guessing a prompt and rerolling. It is a paid call that generates nothing, so use it when the brief is genuinely unclear — not when the user already told you what they want.

Recovering a Timed-Out Call

Both tools are async; on timeout the error carries a task_id and the job keeps running (already charged). Call get_sfx_task(task_id), or get_generation_task(task_id) on the hosted server, later — see task-recovery.

Output Files

  • video_to_sound: a single .wav, named from music_prompt (falling back to sfx_prompt, then sound-<first 8 chars of the task id>).
  • video_to_video_sound: a single .mp4 with the mix muxed in, named the same way (fallback v2v-sound-<first 8 chars of the task id>).

Error Handling

Common errors: 401 invalid key, 402 insufficient balance / trial exhausted, 413 file too large, 422 invalid parameters or malformed segments, 429 rate limit. See the account skill.

© sonilo-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in video-to-sound of sonilo-ai/skills.

Open the folder on GitHubat commit 1ce1bd8

Compare with similar skills

Video To Sound next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video To Sound compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video To Sound this skillsonilo-ai/skills115—~2.6kAutomated safety check: NotesMIT
Musicguaardvark/guaardvark258—~710Automated safety check: PassMIT
Scenario Audioscenario-labs/skills946—~3kAutomated safety check: PassMIT
ShowtimeFavioVazquez/showtime220—~3kAutomated safety check: PassMIT
BlockrunBlockRunAI/blockrun-mcp391—~2.7kAutomated safety check: PassMIT
Videoguaardvark/guaardvark258—~1.2kAutomated safety check: PassMIT

Similar skills

  • Music

    guaardvark/guaardvark

    Generate full songs with vocals or instrumentals (ACE-Step) and sound effects or ambience (Stable Audio Open) on the user's GPU through Guaardvark's Audio Foundry.

    258 GitHub stars~710 tokensUpdated today
    Media & CreativeAuto-check passed
  • Scenario Audio

    scenario-labs/skills

    A skill your agent uses when generating or handling audio on Scenario via MCP.

    946 GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    220 GitHub stars~3k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Blockrun

    BlockRunAI/blockrun-mcp

    Pay-per-call access to AI models, real-time data, media generation and multi-chain RPC over x402 micropayments (USDC on Base or Solana), or a BlockRun account API key.

    391 GitHub stars~2.7k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Video

    guaardvark/guaardvark

    Generate video clips on the user's own GPU through Guaardvark: text-to-video, image-to-video, first+last frame animation, clips with their own soundtrack and dialogue (MiniMax H3), short looping…

    258 GitHub stars~1.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Audio And Video

    glifxyz/glif-mcp-server

    Make or edit audio and video with Glif, from a text brief or from a reference image, video or audio file.

    213 GitHub stars~1.6k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed

More from sonilo-ai/skills

All 13 skills in this repo
  • Text To Sfx

    sonilo-ai/skills

    Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match.

    115 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • Video To Music

    sonilo-ai/skills

    Score a video with original music using Sonilo — the model watches the cut and matches pacing, motion, and emotion, returning either the audio or a new video with the score muxed in.

    115 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Account

    sonilo-ai/skills

    Check the Sonilo account's available services, rate limits, free-trial allowance, and usage/billing history.

    115 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check: notes
  • Audio Ducking

    sonilo-ai/skills

    Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.

    115 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Audio Playback

    sonilo-ai/skills

    Play a local audio file through the system's default speakers using Sonilo's MCP server.

    115 GitHub stars~813 tokensUpdated 2 days ago
    Auto-check: notes
  • Auto Dubbing

    sonilo-ai/skills

    Dub a video into one or more other languages using Sonilo, translating and re-voicing the speech into a new video per language.

    115 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes

Questions about Video To Sound

What does Video To Sound do?

Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo. Video To Sound is an agent skill from sonilo-ai/skills. Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo.

When should I use Video To Sound?

Video To Sound fits situations like: tasks that involve Music and audio generation.

How do I install Video To Sound in Claude Code?

Run `npx skills add sonilo-ai/skills --skill video-to-sound -a claude-code`. Or copy the skill folder (video-to-sound in sonilo-ai/skills) into .claude/skills/video-to-sound in your project. Claude Code loads it when a task matches its description.

How do I install Video To Sound in Codex?

Run `npx skills add sonilo-ai/skills --skill video-to-sound -a codex`. Or copy the skill folder (video-to-sound in sonilo-ai/skills) into .agents/skills/video-to-sound in your project. Codex loads it when a task matches its description.

Can I use Video To Sound in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sonilo-ai/skills --skill video-to-sound -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-to-sound, .gemini/skills/video-to-sound, .github/skills/video-to-sound and .opencode/skills/video-to-sound in your project.

What does Video To Sound need to run?

Going by SKILL.md and its folder, Video To Sound needs the command-line tools its instructions call (pip, npm and curl) and credentials named SONILO_API_KEY. Our summary lists: Python 3; Node.js; A credential in SONILO_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, mcp__sonilo__*. Compatibility (from SKILL.md): Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill..

Does Video To Sound access the network?

SKILL.md names 1 domain. In commands or code: api.sonilo.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Video To Sound safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Video To Sound use?

Video To Sound is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video To Sound use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video To Sound?

Skills that share tags, products or a category with Video To Sound: Music (guaardvark/guaardvark, 258 stars), Scenario Audio (scenario-labs/skills, 946 stars), Showtime (FavioVazquez/showtime, 220 stars) and Blockrun (BlockRunAI/blockrun-mcp, 391 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video To Sound?

sonilo-ai (a GitHub organization) maintains it in sonilo-ai/skills, which has 115 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 9, 2026.

Source: sonilo-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.