Music
guaardvark/guaardvark
Generate full songs with vocals or instrumentals (ACE-Step) and sound effects or ambience (Stable Audio Open) on the user's GPU through Guaardvark's Audio Foundry.
Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo.
$ npx skills add sonilo-ai/skills --skill video-to-sound -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install sonilo-ai/skills video-to-sound --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/video-to-sound .claude/skills/video-to-sound && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-to-sound" agent skill from https://github.com/sonilo-ai/skills/tree/main/video-to-sound into .claude/skills/video-to-sound/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-to-sound", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/sonilo-ai/skills/tree/main/video-to-soundType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add sonilo-ai/skills --skill video-to-sound -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install sonilo-ai/skills video-to-sound --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/video-to-sound .agents/skills/video-to-sound && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-to-sound" agent skill from https://github.com/sonilo-ai/skills/tree/main/video-to-sound into .agents/skills/video-to-sound/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-to-sound", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sonilo-ai/skills --skill video-to-sound -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install sonilo-ai/skills video-to-sound --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/video-to-sound .cursor/skills/video-to-sound && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-to-sound" agent skill from https://github.com/sonilo-ai/skills/tree/main/video-to-sound into .cursor/skills/video-to-sound/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-to-sound", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/sonilo-ai/skills.git --path video-to-sound--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add sonilo-ai/skills --skill video-to-sound -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install sonilo-ai/skills video-to-sound --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/video-to-sound .gemini/skills/video-to-sound && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-to-sound" agent skill from https://github.com/sonilo-ai/skills/tree/main/video-to-sound into .gemini/skills/video-to-sound/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-to-sound", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install sonilo-ai/skills video-to-soundInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add sonilo-ai/skills --skill video-to-sound -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/video-to-sound .github/skills/video-to-sound && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-to-sound" agent skill from https://github.com/sonilo-ai/skills/tree/main/video-to-sound into .github/skills/video-to-sound/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-to-sound", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add sonilo-ai/skills --skill video-to-sound -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install sonilo-ai/skills video-to-sound --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/video-to-sound .opencode/skills/video-to-sound && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-to-sound" agent skill from https://github.com/sonilo-ai/skills/tree/main/video-to-sound into .opencode/skills/video-to-sound/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-to-sound", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-to-soundGenerate music AND sound effects for a video in a single balanced, single-charge call using Sonilo.
Video To Sound is an agent skill from sonilo-ai/skills. Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo. Use instead of calling the music and sound-effects skills separately for the same video — the two layers are mixed and ducked against each other by the backend. Returns a mixed audio track, or a new video with it muxed in.
Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires Sonilo through either transport — the MCP server connected, or the sonilo CLI installed and signed in — plus credentials: a sonilo login sign-in, the…
It sits in Media & Creative, covering Music and audio generation. It works with Model Context Protocol. The repository describes itself as: Agent skills for Sonilo's licensed music, sound-effects, dubbing, and audio-ducking API. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 1ce1bd8. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadWritemcp__sonilo__*From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pipnpmcurlFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.sonilo.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
SONILO_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.
From compatibility in the SKILL.md frontmatter.
Video To Sound loads about 2.6k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,096 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Write, mcp__sonilo__*Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from sonilo-ai/skills at commit 1ce1bd8, republished under its MIT licence (© sonilo-ai). 1,096 words, ~2,628 tokens.
.claude/skills/video-to-sound/SKILL.md (or your agent's skills folder).Generate a music bed and sound effects for a video in one call, balanced against each other and mixed by the backend — one charge instead of two separate generations. Use this whenever a video needs a full soundtrack (score + SFX), not just one or the other.
Setup: See the setup-api-key skill.
⚠️ Cost: makes one API call that may incur charges (billed once, not twice, even though it produces both layers). Only call when explicitly requested.
Pick one at the start of the session and stay on it. Do not mix the two inside a single job, and do not announce the choice.
video_to_sound and friends) — use them. This is the preferred path: it needs no shell, and it is the only one that survives a very long generation. If a call fails to authenticate — rather than failing on its inputs — this transport is not usable in this session: go to 2 instead of retrying it.sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe with sonilo account, not sonilo whoami: whoami exits 0 even when signed out, so it cannot tell the two states apart.api.sonilo.com with curl to work around it; both transports handle uploads, polling and retries that a bare request does not.video_to_sound(
video_path="~/Desktop/trailer.mp4",
music_prompt="Cinematic, building tension",
sfx_prompt="Footsteps, wind, distant thunder"
)video_to_video_sound(
video_path="~/Desktop/trailer.mp4",
music_prompt="Cinematic, building tension"
)pip install sonilo)from sonilo import Sonilo
client = Sonilo() # reads SONILO_API_KEY
mix = client.video_to_sound.generate(
video="trailer.mp4",
music_prompt="Cinematic, building tension",
sfx_prompt="Footsteps, wind, distant thunder",
)
mix.save("soundtrack.wav")
video = client.video_to_video_sound.generate(video="trailer.mp4", music_prompt="Cinematic, building tension")
video.save("scored.mp4")npm install sonilo)import { SoniloClient, download } from "sonilo";
import { writeFile } from "node:fs/promises";
const client = new SoniloClient(); // reads SONILO_API_KEY
const mix = await client.videoToSound.generate({
video: "./trailer.mp4",
musicPrompt: "Cinematic, building tension",
sfxPrompt: "Footsteps, wind, distant thunder",
});
await writeFile("soundtrack.wav", await download(mix.output_url));
const video = await client.videoToVideoSound.generate({
video: "./trailer.mp4",
musicPrompt: "Cinematic, building tension",
});
await writeFile("scored.mp4", await download(video.output_url));npm install -g sonilo-cli or pip install sonilo-cli)sonilo video-to-sound --video trailer.mp4 \
--music-prompt "Cinematic, building tension" --sfx-prompt "Footsteps, wind, distant thunder" \
--output soundtrack.wav
sonilo video-to-video-sound --video trailer.mp4 --music-prompt "Cinematic, building tension"Unlike the music/sound-effects skills, both tools here have CLI commands. --stem music/--stem sfx (repeatable) additionally saves the individual layers next to the combined output.
curl -X POST "https://api.sonilo.com/v1/video-to-sound" \
-H "Authorization: Bearer $SONILO_API_KEY" \
-F "video=@trailer.mp4" \
-F "music_prompt=Cinematic, building tension" \
-F "sfx_prompt=Footsteps, wind, distant thunder"
# -> {"task_id": "..."} poll GET /v1/tasks/{task_id}Both endpoints are task-based (202 + poll), same as the sound-effects tools — the MCP tool waits for you.
| Tool | Description |
|---|---|
video_to_sound(video_path? | video_url?, music_prompt?, sfx_prompt?, segments?, preserve_speech?, ducking?, output_format?, variants_num?, output_directory?) | Generate and mix music + SFX for a video, returns a single audio file. |
video_to_video_sound(video_path? | video_url?, music_prompt?, sfx_prompt?, segments?, keep_original_sound?, preserve_speech?, ducking?, variants_num?, output_directory?) | Same, but returns a new .mp4 with the mixed soundtrack muxed in. By default the source's own audio is dropped — see keep_original_sound. |
| Parameter | Type | Default | Notes |
|---|---|---|---|
video_path | string | — | .mp4/.mov/.webm/.m4v/.gif (gif must be animated). Max 480s (8 min), subject to the account's upload-size cap. |
video_url | string | — | HTTPS/HTTP URL. Exactly one of video_path/video_url. |
music_prompt | string | — | Style hint for the music bed (max 2000 chars). Optional — omit to let Sonilo decide. |
sfx_prompt | string | — | Description of the SFX layered over the music (max 2000 chars). Optional. |
segments | list[dict] | — | Per-segment SFX descriptions — same schema and validation rules as in the video-to-sfx skill. Max 30 segments. |
preserve_speech | bool | false | Keep the source video's speech audible in the mix. |
ducking | bool | false | Brings the source video's own speech into the mix and dips the generated music under it. Off by default: with ducking and preserve_speech both unset, the result carries the generated music and effects alone and no music_processed stem exists. Pass true for any video with dialogue or narration that should stay audible. |
keep_original_sound | bool | false | video_to_video_sound only. Keeps the whole source track (dialogue, room tone, existing effects) with the generated mix over it, rather than replacing it. Add ducking=true to dip the mix under the voice instead of a flat blend. Supersedes preserve_speech. |
output_format | string | wav | video_to_sound only — video_to_video_sound always returns an .mp4. wav, m4a, or mp3 (320 kbps). Sets the combined track's container only; stems keep their own native formats. |
variants_num | int | 1 | 1–10 distinct mixes in one request, one file each. Cost scales linearly and any value above 1 is never free-trial covered — confirm the count with the user before calling. |
output_directory | string | SONILO_MCP_BASE_PATH | Absolute, or relative to the base path. |
No prompt is required — the model reads the cut. A short structured brief adds your intent on top. Since this endpoint generates music and SFX in one balanced call, both crafts apply:
video_to_music + video_to_sfx. The two layers are balanced against each other by the backend (so the SFX doesn't fight the score), and it's one charge, not two.music_prompt and sfx_prompt are optional — you can leave both unset and let Sonilo interpret the whole scene, or set just one to steer that layer while leaving the other automatic.ducking is off by default — turn it on for anything with a voice. Left off, the source speech is not in the mix at all: the output is generated music and effects only. That is the right default for a silent or music-only clip and the wrong one for a talking head, so check the source audio before calling (see the pre-flight reference) rather than after the user tells you the narration is gone.video_to_video_sound, the source audio is dropped unless you say otherwise. keep_original_sound=true keeps the whole original track under the generated mix; preserve_speech=true keeps only the isolated speech. If a user reports "my dialogue disappeared", this is the fix.video_to_video_sound instead of video_to_sound.sfx_segments + sfx_prompt) read off the footage, which beats guessing a prompt and rerolling. It is a paid call that generates nothing, so use it when the brief is genuinely unclear — not when the user already told you what they want.Both tools are async; on timeout the error carries a task_id and the job keeps running (already charged). Call get_sfx_task(task_id), or get_generation_task(task_id) on the hosted server, later — see task-recovery.
video_to_sound: a single .wav, named from music_prompt (falling back to sfx_prompt, then sound-<first 8 chars of the task id>).video_to_video_sound: a single .mp4 with the mix muxed in, named the same way (fallback v2v-sound-<first 8 chars of the task id>).Common errors: 401 invalid key, 402 insufficient balance / trial exhausted, 413 file too large, 422 invalid parameters or malformed segments, 429 rate limit. See the account skill.
© sonilo-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in video-to-sound of sonilo-ai/skills.
Open the folder on GitHubat commit 1ce1bd8
Video To Sound next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Video To Sound this skillsonilo-ai/skills | 115 | — | ~2.6k | Automated safety check: Notes | MIT | |
| Musicguaardvark/guaardvark | 258 | — | ~710 | Automated safety check: Pass | MIT | |
| Scenario Audioscenario-labs/skills | 946 | — | ~3k | Automated safety check: Pass | MIT | |
| ShowtimeFavioVazquez/showtime | 220 | — | ~3k | Automated safety check: Pass | MIT | |
| BlockrunBlockRunAI/blockrun-mcp | 391 | — | ~2.7k | Automated safety check: Pass | MIT | |
| Videoguaardvark/guaardvark | 258 | — | ~1.2k | Automated safety check: Pass | MIT |
guaardvark/guaardvark
Generate full songs with vocals or instrumentals (ACE-Step) and sound effects or ambience (Stable Audio Open) on the user's GPU through Guaardvark's Audio Foundry.
scenario-labs/skills
A skill your agent uses when generating or handling audio on Scenario via MCP.
FavioVazquez/showtime
A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…
BlockRunAI/blockrun-mcp
Pay-per-call access to AI models, real-time data, media generation and multi-chain RPC over x402 micropayments (USDC on Base or Solana), or a BlockRun account API key.
guaardvark/guaardvark
Generate video clips on the user's own GPU through Guaardvark: text-to-video, image-to-video, first+last frame animation, clips with their own soundtrack and dialogue (MiniMax H3), short looping…
glifxyz/glif-mcp-server
Make or edit audio and video with Glif, from a text brief or from a reference image, video or audio file.
sonilo-ai/skills
Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match.
sonilo-ai/skills
Score a video with original music using Sonilo — the model watches the cut and matches pacing, motion, and emotion, returning either the audio or a new video with the score muxed in.
sonilo-ai/skills
Check the Sonilo account's available services, rate limits, free-trial allowance, and usage/billing history.
sonilo-ai/skills
Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.
sonilo-ai/skills
Play a local audio file through the system's default speakers using Sonilo's MCP server.
sonilo-ai/skills
Dub a video into one or more other languages using Sonilo, translating and re-voicing the speech into a new video per language.
Works with
Categories
Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo. Video To Sound is an agent skill from sonilo-ai/skills. Generate music AND sound effects for a video in a single balanced, single-charge call using Sonilo.
Video To Sound fits situations like: tasks that involve Music and audio generation.
Run `npx skills add sonilo-ai/skills --skill video-to-sound -a claude-code`. Or copy the skill folder (video-to-sound in sonilo-ai/skills) into .claude/skills/video-to-sound in your project. Claude Code loads it when a task matches its description.
Run `npx skills add sonilo-ai/skills --skill video-to-sound -a codex`. Or copy the skill folder (video-to-sound in sonilo-ai/skills) into .agents/skills/video-to-sound in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sonilo-ai/skills --skill video-to-sound -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-to-sound, .gemini/skills/video-to-sound, .github/skills/video-to-sound and .opencode/skills/video-to-sound in your project.
Going by SKILL.md and its folder, Video To Sound needs the command-line tools its instructions call (pip, npm and curl) and credentials named SONILO_API_KEY. Our summary lists: Python 3; Node.js; A credential in SONILO_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, mcp__sonilo__*. Compatibility (from SKILL.md): Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill..
SKILL.md names 1 domain. In commands or code: api.sonilo.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.
Video To Sound is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Video To Sound: Music (guaardvark/guaardvark, 258 stars), Scenario Audio (scenario-labs/skills, 946 stars), Showtime (FavioVazquez/showtime, 220 stars) and Blockrun (BlockRunAI/blockrun-mcp, 391 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
sonilo-ai (a GitHub organization) maintains it in sonilo-ai/skills, which has 115 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 9, 2026.
Source: sonilo-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.