Agent skill

Text To Music

by sonilo-ai in sonilo-ai/skills

Generate music from a text prompt using Sonilo — instrumental tracks, background beds, jingles, loops — when there is no video to score.

MITAuto-check: notesMedia & Creative

Install Text To Music

skills CLI
$ npx skills add sonilo-ai/skills --skill text-to-music -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sonilo-ai/skills text-to-music --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sonilo-ai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/text-to-music .claude/skills/text-to-music && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
text-to-music
GitHub stars
115
Token cost
~2.7k tokens
SKILL.md length
1,182 words
Files
1
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Generate music from a text prompt using Sonilo — instrumental tracks, background beds, jingles, loops — when there is no video to score.

  • Works in 3 steps: Sonilo MCP tools visible in this session… → No usable Sonilo MCP tools, but sonilo… → Neither — stop and run the setup-api-key…
  • The user describes the music they want in words
  • SKILL.md covers Transport: MCP or CLI, Quick Start, Tool and Parameters, plus 6 more sections
  • Calls curl, pip and npm; reaches api.sonilo.com; needs SONILO_API_KEY

What it does

Text To Music is an agent skill from sonilo-ai/skills. Generate music from a text prompt using Sonilo — instrumental tracks, background beds, jingles, loops — when there is no video to score. Use when the user describes the music they want in words, with or without a length. Every track is licensed and cleared for commercial use. For scoring an existing video, use the video-to-music skill instead.

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires Sonilo through either transport — the MCP server connected, or the sonilo CLI installed and signed in — plus credentials: a sonilo login sign-in, the…

It sits in Media & Creative, covering Music and audio generation and MCP servers. It works with Model Context Protocol. The repository describes itself as: Agent skills for Sonilo's licensed music, sound-effects, dubbing, and audio-ducking API. The licence is MIT.

When your agent uses it

  • The user describes the music they want in words
  • Without a length

Example prompts

  • “/text-to-music”

Requirements

  • Python 3
  • Node.js
  • A credential in SONILO_API_KEY
  • Compatibility (from SKILL.md): Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.
  • Pre-approved tools (allowed-tools): Bash, Read, Write, mcp__sonilo__*

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Sonilo MCP tools visible in this session (text_to_music and friends) — use them. This is the preferred path: it needs no shell, and it is…
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe…
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads…

What it can do on your machine

Read from SKILL.md and the folder at commit 1ce1bd8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write
    • mcp__sonilo__*

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • pip
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.sonilo.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • SONILO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.

    From compatibility in the SKILL.md frontmatter.

Context cost

Text To Music loads about 2.7k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 1,182 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write, mcp__sonilo__*

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sonilo-ai/skills at commit 1ce1bd8, republished under its MIT licence (© sonilo-ai). 1,182 words, ~2,671 tokens.

Download SKILL.mdSave it as .claude/skills/text-to-music/SKILL.md (or your agent's skills folder).
name
text-to-music
description
Generate music from a text prompt using Sonilo — instrumental tracks, background beds, jingles, loops — when there is no video to score. Use when the user describes the music they want in words, with or without a length. Every track is licensed and cleared for commercial use. For scoring an existing video, use the video-to-music skill instead.
allowed-tools
Bash, Read, Write, mcp__sonilo__*
compatibility
Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill.
license
MIT

Sonilo Text-to-Music

Generate music from a text description alone — no video involved. The prompt IS the input, so the brief carries everything: genre, mood, energy arc, instrumentation, and what must not appear. Every track is licensed (music licensed via Shutterstock) and cleared for commercial use on social, brand content, and advertising.

Setup: See the setup-api-key skill to connect the Sonilo MCP server and authenticate — sonilo login (no key) or SONILO_API_KEY.

⚠️ Cost: this tool makes an API call that may incur charges. Only call it when the user has actually asked for a generation. Check get_account_services (see the account skill) if you're unsure whether free-trial runs remain.

Scoring a video instead? Use video-to-music — it matches pacing, motion, and emotion to the actual cut, which a text prompt cannot do.

Transport: MCP or CLI

Pick one at the start of the session and stay on it. Do not mix the two inside a single job, and do not announce the choice.

  1. Sonilo MCP tools visible in this session (text_to_music and friends) — use them. This is the preferred path: it needs no shell, and it is the only one that survives a very long generation. If a call fails to authenticate — rather than failing on its inputs — this transport is not usable in this session: go to 2 instead of retrying it.
  2. No usable Sonilo MCP tools, but sonilo account exits 0 — use the CLI commands below. Same API, same account, same credential file. Probe with sonilo account, not sonilo whoami: whoami exits 0 even when signed out, so it cannot tell the two states apart.
  3. Neither — stop and run the setup-api-key skill. Do not call api.sonilo.com with curl to work around it; both transports handle uploads, polling and retries that a bare request does not.

Quick Start

text_to_music(
    prompt="A chill lo-fi hip hop beat with jazzy piano chords",
    duration=30
)

Saves the generated file(s) to SONILO_MCP_BASE_PATH (~/Desktop by default) and returns the saved path(s) as text.

Python (pip install sonilo)
python
from sonilo import Sonilo

client = Sonilo()  # reads SONILO_API_KEY

track = client.text_to_music.generate(prompt="A chill lo-fi hip hop beat with jazzy piano chords", duration=30)
track.save("output.mp3")
JavaScript / TypeScript (npm install sonilo)
ts
import { SoniloClient } from "sonilo";

const client = new SoniloClient(); // reads SONILO_API_KEY

const track = await client.textToMusic.generate({
  prompt: "A chill lo-fi hip hop beat with jazzy piano chords",
  duration: 30,
});
CLI (npm install -g sonilo-cli or pip install sonilo-cli)
bash
sonilo text-to-music --prompt "A chill lo-fi hip hop beat with jazzy piano chords" --duration 30
cURL (raw REST API, no MCP host)
bash
curl -X POST "https://api.sonilo.com/v1/text-to-music" \
  -H "Authorization: Bearer $SONILO_API_KEY" \
  --data-urlencode "prompt=A chill lo-fi hip hop beat with jazzy piano chords" \
  --data-urlencode "duration=30" \
  --output output.m4a

Tool

ToolDescription
text_to_music(prompt, duration?, output_format?, variants_num?, stems?, output_directory?)Generate music from a text description only — no video.

Parameters

ParameterTypeDefaultNotes
promptstring—Required. 1–1000 chars.
durationintinferred from the promptOptional, 5–360 seconds. Omit it when the user names no length — Sonilo reads one out of the prompt, so a "short jingle" comes back short and a "full-length closing-credits piece" comes back long. A vague prompt infers a long track ("lofi" alone resolves to about 180 seconds) and you are billed for it, so ask the user for a length when the cost matters.
variants_numint11–10. Generates that many distinct creative directions in one request — different takes, not re-renders of one. Cost scales linearly with the count, and any value above 1 is never covered by the free trial, so confirm the number with the user first. Above 1 writes one file per variant and forces the backend's async mode.
output_formatstringm4am4a or wav. wav triggers the backend's async mode internally — no user-facing "mode" param needed.
stemsboolfalseFree. Additionally splits each generated track into four separated instrument tracks — drums, bass, vocals, other — returned alongside the untouched full mix. Async-only on REST (stems=true without mode=async is a 400). See Stems.
output_directorystringSONILO_MCP_BASE_PATHAbsolute, or relative to the base path.

Stems

stems=true additionally returns each generated track split into four separated instrument tracks — drums, bass, vocals, other — free of charge. The full mix is untouched; the stems arrive alongside it in the task result as a stems array next to audio:

json
"stems": [
  {
    "stream_index": 0,
    "drums":  { "url": "…", "content_type": "audio/mp4", "file_size": 2913044 },
    "bass":   { "url": "…", "content_type": "audio/mp4", "file_size": 2870211 },
    "vocals": { "url": "…", "content_type": "audio/mp4", "file_size": 2794560 },
    "other":  { "url": "…", "content_type": "audio/mp4", "file_size": 3011830 }
  }
]

What matters when you use it:

  • Async only on REST. stems=true requires mode=async (a 400 otherwise): you get a 202 + task_id and poll /v1/tasks/{task_id}. The MCP tools are always async, so on the hosted server the param just works.
  • Available on every surface (verified 2026-08-17): REST, the hosted MCP server, the local sonilo-mcp package (>= 0.18.0), the SDKs (sonilo npm >= 0.16.0, PyPI >= 0.15.0), and the CLIs (--stems, npm sonilo-cli >= 0.15.0, PyPI sonilo-cli >= 0.14.0).
  • Match stems to tracks by stream_index, never by array position. A stream whose separation failed is simply absent, so stems can be shorter than audio.
  • stems_error is not a failed generation. When separation failed wholly or partly, or was skipped, the task carries a stems_error string — possibly alongside a partial stems array. The generation itself succeeded and every audio URL is valid: treat missing stems as a missing extra, never as a reason to retry or refund.
  • Timing: separation runs after generation finishes — typically another 2–6 min, giving up after 30 min (then stems_error).
  • The four stem names are fixed (htdemucs separation): melodic instruments — piano, synths, guitar, strings — land in other, and on instrumental tracks vocals is near-silent. That is correct behavior, not a bug.
  • Formats: stems normally follow output_format; trust each stem's content_type for what was actually delivered.
bash
# REST: submit with stems, then poll the task
curl -X POST "https://api.sonilo.com/v1/text-to-music" \
  -H "Authorization: Bearer $SONILO_API_KEY" \
  --data-urlencode "prompt=A chill lo-fi hip hop beat with jazzy piano chords" \
  --data-urlencode "duration=30" \
  --data-urlencode "mode=async" \
  --data-urlencode "stems=true"
# → {"task_id": "…"} — poll GET /v1/tasks/{task_id} for audio + stems
Show full SKILL.md (368 more words)Show less

Prompting

The prompt is the only input — there is no footage to lean on. Describe genre, mood, tempo, instrumentation, and the energy arc; "A driving synthwave track with arpeggiated leads" beats "electronic music." Describe structure in words too ("builds for 10 s, drops, outro"), since there is no cut to infer it from. Name exclusions explicitly — the sounds that must not appear.

The same craft vocabulary applies as for video scoring, minus the video pre-flight: references/music-prompting.md.

Generate once and iterate on the prompt, not on rerolls — failed runs auto-refund, but your own retry is a new charge.

Workflow Tips

  • If the user has a finished video, you are in the wrong skill. video-to-music syncs to the actual cut instead of producing a generic track of matching length.
  • Duration is optional, and a guess is worse than leaving it out. Pass the length the user named. If they named none, omit duration and Sonilo reads one out of the prompt — but a vague prompt resolves long ("lofi" alone is about 180 s) and is billed at that length, so ask first when cost matters.
  • Several takes in one go: variants_num=3 returns three distinct directions for one request instead of three re-rolls. It costs 3×, and it is never free-trial covered — say the price before calling.
  • User wants the track's instruments as separate files (to remix, re-balance, or drop one)? stems=true — it's free, but async-only and not on every surface yet; see Stems.
  • Content restriction: prompts cannot reference specific artists, bands, or copyrighted lyrics.

Recovering a Timed-Out Call

text_to_music streams its result in one call unless output_format="wav", variants_num above 1, or stems=true triggers the backend's async mode. If an async variant times out, the error message includes a task_id; the generation keeps running (and is already charged) on the backend. Call get_sfx_task(task_id) — get_generation_task(task_id) on the hosted server — to retrieve the result; see the task-recovery skill.

Output Files

.m4a by default (.wav if requested), named from the prompt (slugified) or sonilo-<timestamp>.m4a. Multiple parallel streams get a -<index> suffix.

Error Handling

Common errors: 401 invalid key, 402 insufficient balance / trial exhausted, 422 invalid parameters (e.g. duration out of range), 429 rate limit. See the account skill to check trial/usage before a call.

© sonilo-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in text-to-music of sonilo-ai/skills.

Open the folder on GitHubat commit 1ce1bd8

Compare with similar skills

Text To Music next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Text To Music compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Text To Music this skillsonilo-ai/skills115—~2.7kAutomated safety check: NotesMIT
Resolve Audiosamuelgursky/davinci-resolve-mcp3.5k—~1.3kAutomated safety check: PassMIT
Musicguaardvark/guaardvark258—~710Automated safety check: PassMIT
Fal Assetsrehan-remade/universal-modder6.5k—~2kAutomated safety check: NotesMIT
Scenario Audioscenario-labs/skills946—~3kAutomated safety check: PassMIT
OpenStoryline Install HelperFireRedTeam/FireRed-OpenStoryline3.5k—~1.5kAutomated safety check: NotesApache-2.0

Similar skills

  • Resolve Audio

    samuelgursky/davinci-resolve-mcp

    Audio and Fairlight work in the DaVinci Resolve MCP. An agent skill from samuelgursky/davinci-resolve-mcp.

    3.5k GitHub stars~1.3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Music

    guaardvark/guaardvark

    Generate full songs with vocals or instrumentals (ACE-Step) and sound effects or ambience (Stable Audio Open) on the user's GPU through Guaardvark's Audio Foundry.

    258 GitHub stars~710 tokensUpdated today
    Media & CreativeAuto-check passed
  • Fal Assets

    rehan-remade/universal-modder

    Generate game assets with fal (fal.ai) through the fal MCP server, the um fal CLI (REST) or fal api.

    6.5k GitHub stars~2k tokensUpdated today
    Game DevelopmentAuto-check: notes
  • Scenario Audio

    scenario-labs/skills

    A skill your agent uses when generating or handling audio on Scenario via MCP.

    946 GitHub stars~3k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • OpenStoryline Install Helper

    FireRedTeam/FireRed-OpenStoryline

    Installs, repairs and starts a local source checkout of FireRed-OpenStoryline, from prerequisites and a venv to resources, config and the MCP and web servers.

    3.5k GitHub stars~1.5k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Free AI-first short-form video publishing to TikTok, Instagram Reels, YouTube Shorts, X, and Facebook from AI agents through Taisly.

    217 GitHub starsUsed in 1 repo~1.5k tokens
    Media & CreativeAuto-check passed

More from sonilo-ai/skills

All 13 skills in this repo
  • Text To Sfx

    sonilo-ai/skills

    Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match.

    115 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • Video To Music

    sonilo-ai/skills

    Score a video with original music using Sonilo — the model watches the cut and matches pacing, motion, and emotion, returning either the audio or a new video with the score muxed in.

    115 GitHub starsUsed in 1 repo~4k tokens
    Auto-check: notes
  • Account

    sonilo-ai/skills

    Check the Sonilo account's available services, rate limits, free-trial allowance, and usage/billing history.

    115 GitHub stars~1.4k tokensUpdated 2 days ago
    Auto-check: notes
  • Audio Ducking

    sonilo-ai/skills

    Duck a music bed under a voice track using Sonilo — automatically lowers the music wherever the voice speaks and lifts it back in the gaps.

    115 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check: notes
  • Audio Playback

    sonilo-ai/skills

    Play a local audio file through the system's default speakers using Sonilo's MCP server.

    115 GitHub stars~813 tokensUpdated 2 days ago
    Auto-check: notes
  • Auto Dubbing

    sonilo-ai/skills

    Dub a video into one or more other languages using Sonilo, translating and re-voicing the speech into a new video per language.

    115 GitHub stars~4.5k tokensUpdated 2 days ago
    Auto-check: notes

Questions about Text To Music

What does Text To Music do?

Generate music from a text prompt using Sonilo — instrumental tracks, background beds, jingles, loops — when there is no video to score. Text To Music is an agent skill from sonilo-ai/skills. Generate music from a text prompt using Sonilo — instrumental tracks, background beds, jingles, loops — when there is no video to score.

When should I use Text To Music?

Text To Music fits situations like: the user describes the music they want in words; without a length.

How do I install Text To Music in Claude Code?

Run `npx skills add sonilo-ai/skills --skill text-to-music -a claude-code`. Or copy the skill folder (text-to-music in sonilo-ai/skills) into .claude/skills/text-to-music in your project. Claude Code loads it when a task matches its description.

How do I install Text To Music in Codex?

Run `npx skills add sonilo-ai/skills --skill text-to-music -a codex`. Or copy the skill folder (text-to-music in sonilo-ai/skills) into .agents/skills/text-to-music in your project. Codex loads it when a task matches its description.

Can I use Text To Music in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sonilo-ai/skills --skill text-to-music -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/text-to-music, .gemini/skills/text-to-music, .github/skills/text-to-music and .opencode/skills/text-to-music in your project.

What does Text To Music need to run?

Going by SKILL.md and its folder, Text To Music needs the command-line tools its instructions call (curl, pip and npm) and credentials named SONILO_API_KEY. Our summary lists: Python 3; Node.js; A credential in SONILO_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write, mcp__sonilo__*. Compatibility (from SKILL.md): Requires Sonilo through either transport — the MCP server connected, or the `sonilo` CLI installed and signed in — plus credentials: a `sonilo login` sign-in, the hosted OAuth plugin, or SONILO_API_KEY. See the setup-api-key skill..

Does Text To Music access the network?

SKILL.md names 1 domain. In commands or code: api.sonilo.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Text To Music safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Text To Music use?

Text To Music is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Text To Music use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Text To Music?

Skills that share tags, products or a category with Text To Music: Resolve Audio (samuelgursky/davinci-resolve-mcp, 3.5k stars), Music (guaardvark/guaardvark, 258 stars), Fal Assets (rehan-remade/universal-modder, 6.5k stars) and Scenario Audio (scenario-labs/skills, 946 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Text To Music?

sonilo-ai (a GitHub organization) maintains it in sonilo-ai/skills, which has 115 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on October 9, 2026.

Source: sonilo-ai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.