Agent skill

Scenario Audio

by scenario-labs in scenario-labs/skills

A skill your agent uses when generating or handling audio on Scenario via MCP.

MITAuto-check passedMedia & Creative

Install Scenario Audio

skills CLI
$ npx skills add scenario-labs/skills --skill scenario-audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install scenario-labs/skills scenario-audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/scenario-labs/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/scenario-audio .claude/skills/scenario-audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scenario-audio
GitHub stars
946
Token cost
~3k tokens
SKILL.md length
1,550 words
Files
1
Skills in repo
146
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when generating or handling audio on Scenario via MCP.

  • Works in 6 steps: recommend with capability="txt2audio"… → model_schema_get with that model_id.… → model_run with the same model_id and… → …
  • Handling audio on Scenario via MCP
  • SKILL.md covers Overview, Quick reference, What the audio surface covers and Worked example: a game sound…, plus 4 more sections
  • Calls curl and npx

What it does

Scenario Audio is an agent skill from scenario-labs/skills. Use when generating or handling audio on Scenario via MCP. Triggers include music tracks, full-length songs with vocals written from lyrics, background scores, soundtracks, game sound effects, SFX, foley, ambience, looping audio, voiceover, narration, speech, TTS, text-to-speech, dialogue, voice cloning, re-voicing a recording, scoring or adding sound to a video, transcription, or requests to create, wait on, play, or download audio files (MP3, WAV) with Scenario tools.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice and Music and audio generation. It works with Model Context Protocol. The repository describes itself as: Get production-ready images, video, audio, and 3D from any AI agent: skills that pick the right model, price before spending, and keep characters and brands consistent through… The licence is MIT.

When your agent uses it

  • Handling audio on Scenario via MCP
  • Include music tracks
  • Full-length songs with vocals written from lyrics
  • Background scores

Example prompts

  • “/scenario-audio”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. recommend with capability="txt2audio" and the user's own words as prompt ("a game sound effect: a heavy wooden chest creaking open"). The…
  2. model_schema_get with that model_id. Returns the exact fields: prompt plus controls such as duration or looping.
  3. model_run with the same model_id and parameters={"prompt": "heavy wooden treasure chest creaking open, single event, dry, no music"}.
  4. jobs_wait job_ids=["job_xxx"] on any job_id returned without assets (in_progress after a timed-out wait, the backend's queued or…
  5. asset_display asset_id="asset_xxx" to play it inline.
  6. asset_download with no format, then save the returned URL with curl -L.

What it can do on your machine

Read from SKILL.md and the folder at commit f6f8ab7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl
    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and npx, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scenario Audio loads about 3k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 1,550 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from scenario-labs/skills at commit f6f8ab7, republished under its MIT licence (© scenario-labs). 1,550 words, ~2,995 tokens.

Download SKILL.mdSave it as .claude/skills/scenario-audio/SKILL.md (or your agent's skills folder).
name
scenario-audio
description
Use when generating or handling audio on Scenario via MCP. Triggers include music tracks, full-length songs with vocals written from lyrics, background scores, soundtracks, game sound effects, SFX, foley, ambience, looping audio, voiceover, narration, speech, TTS, text-to-speech, dialogue, voice cloning, re-voicing a recording, scoring or adding sound to a video, transcription, or requests to create, wait on, play, or download audio files (MP3, WAV) with Scenario tools.
license
MIT

Scenario Audio Generation

Overview

Scenario generates audio through the same loop as images. The live catalog covers three generation lanes (music, sound effects, voice/speech) plus video-to-audio soundtrack models and audio utilities. Connection and the core generation loop: see the scenario skill. If a sibling skill named here is missing from your available skills, ask the user to install it (npx skills add scenario-labs/skills --skill <name>); unattended, proceed from tool schemas and flag the gap.

Quick reference

StepToolNotes
Find a modelrecommend with the need in the user's own words; search only for a member known by namecapability="txt2audio" covers music, SFX, and TTS; optional, inferred from the prompt when omitted
Inspect inputsmodel_schema_getaudio schemas vary widely: durations, lyrics, voices, looping
Generatemodel_runschema-conformant parameters; wait=false for long jobs
Waitjobs_waitblocks server-side; on timeout re-call with pending_job_ids
Listenasset_displayrenders an inline audio player
Saveasset_downloadreturns a download URL: curl -L -o out.mp3 "<url>"

Find existing audio assets with search target="assets", filters={kind: "audio"}. Team and project scope (team_id, project_id): see the scenario skill.

What the audio surface covers

  • Music: text-to-music models produce short beds or full-length songs with vocals; the song lane has its own contract, below.
  • Sound effects: text-to-SFX models generate short clips from a description; some support seamless looping.
  • Voice and speech: text-to-speech with preset voices, multilingual output, and emotion or pacing controls; some clone a voice from a short clip, and speech-to-speech re-voices a recording.
  • Video to audio: models that score a silent video or add synchronized effects.
  • Utilities: model_scenario-audio-cut, model_scenario-audio-split, model_scenario-audio-extract, and model_scenario-compose-video (fixed ids: each is Scenario's single deterministic tool for its operation, so discovery would only re-derive them); the compositor lays a finished track (score, voiceover, re-voiced take) over a clip as an audio layer, per scenario-video-assembly; for speech-to-text transcription, recommend with the need in the user's own words.
  • Stem separation: one named stem per run (discover with recommend), vocals included, with no instrumental option. Voice isolation returns the clean speech and never the removed music and effects as a second stem, so a two-stem split (voice against everything else) is a gap to report, not a member to keep hunting for.

Per-family contracts: scenario-elevenlabs (speech, dubbing, re-voicing, music, SFX), scenario-ace-step and scenario-minimax-music (songs), scenario-sonilo (SFX and video scoring).

Worked example: a game sound effect

  1. recommend with capability="txt2audio" and the user's own words as prompt ("a game sound effect: a heavy wooden chest creaking open"). The ranking returns txt2audio models such as model_elevenlabs-sound-effects-v2 (example only).
  2. model_schema_get with that model_id. Returns the exact fields: prompt plus controls such as duration or looping.
  3. model_run with the same model_id and parameters={"prompt": "heavy wooden treasure chest creaking open, single event, dry, no music"}.
  4. jobs_wait job_ids=["job_xxx"] on any job_id returned without assets (in_progress after a timed-out wait, the backend's queued or in-progress after wait=false), re-calling with the returned pending_job_ids on timeout.
  5. asset_display asset_id="asset_xxx" to play it inline.
  6. asset_download with no format, then save the returned URL with curl -L.

Prompting tips:

  • SFX: name the source, material, action, and acoustic space, and say what to exclude ("no music", "no reverb"). One event per clip; generate variations as separate runs.
  • Music: give genre, mood, tempo, and instrumentation. Short beds usually take a single prompt, with duration or looping in the schema.
  • Speech: keep the text field to the words to speak; voice, language, emotion, and pacing live in separate schema fields or inline tags.

Speech and dialogue

Discover a voice member with recommend (capability="txt2audio", the user's words, "two-person dialogue" when it is one), then read its text field's description in model_schema_get: that is where a member usually states its own delivery grammar.

  • Tag syntax is per member, even within one family: square brackets ([whispers]), angle brackets (<sigh>, <short pause>), parentheses ((sighs)), or wrapping pairs (<whisper>text</whisper>), and some read none. A tag in the wrong grammar can be spoken aloud, so copy the spelling from the text field's description; where it names none, from the member's catalog description or recommend's notes, and with no source write no tags. A correctly spelled tag is a request, not a guarantee: two identical runs have disagreed on one, so have the user listen before a take is final and re-run a take whose tag was skipped.
  • Two voices, one take. A member with a multi-speaker array (up to 2 rows of a speaker label and a voice at authoring time) reads the text as turns, one Name: line per turn, each Name matching a row's label exactly; an unprefixed line continues the previous turn, and the single-voice field is ignored in that mode. Preset voice names say nothing about gender, age, or tone, so state which preset plays whom and let the user confirm before the paid run. A scene with more speakers is split into takes of at most two, delivered in order or laid over the picture per scenario-video-assembly.
  • Text is capped and priced. The text field carries a max_length (5000 characters on one member at authoring time), an overrun is a 400, and cost_impact marks it as the price driver: split a long script at turn boundaries and dry_run the first take.
  • Pin the language through the schema's language field when there is one, rather than naming it in the text.
Show full SKILL.md (671 more words)Show less

Songs with vocals

A full-length song is not a longer music bed, and song schemas vary more than the rest of the lane, so model_schema_get decides the shape: a style prompt plus a separate lyric sheet, one prose prompt carrying both, or an ordered section array with per-section text and styles.

  • Words never go in a style field. Where the schema splits the two, the style field carries genre, mood, tempo, key, vocal style, and instrumentation; the lyric field carries the words, shaped by section tags such as [Verse] and [Chorus].
  • Instrumental and auto-lyrics are flags where the schema has them; asking for either in prose is unreliable, and where no flag exists the text fields are the only lever. Flipping the instrumental flag on a rerun gives a different take, not the same song: seed, where a model has one, only repeats identical settings. For an instrumental of a track you already have, try an audio2audio cover model (discover with recommend), checking the schema since not all carry the flag.
  • Text fields are length-capped per model and field, and going over is a 400 rather than a truncation.

Where the schema exposes a duration field (flagged cost_impact), it caps both length and price; where none exists, the lyric sheet or prompt sets both. Either way, price the song with dry_run: true before committing, then launch with wait: false; both are model_run arguments, not parameters keys.

Repeatability and batching are per member, not per lane: at authoring time the repaint members took numOutputs (1 to 4) and no seed, so a repaint that must keep the same singer is run as a batch and picked from, while the section composers had seed and no numOutputs. Read both off model_schema_get before promising either.

Extending a song

Making an existing track longer is its own audio2audio lane, not a longer text-to-music run: recommend with capability="audio2audio" and the extension need in the user's words; use search only when the member is already known by name. The member that does it takes the song as audio and an ordered sections array (up to 30 at authoring time) where each entry either keeps a slice of the original (sourceStartSeconds and sourceEndSeconds) or generates a new one (text with [Verse]-style tags, durationSeconds, positiveStyles as an array), with contextAdherence deciding how closely new sections follow their neighbors. Kept slices bill like generated audio of the same length, so dry_run the whole plan first. The repaint, edit and add-layer members regenerate inside the original's duration and never lengthen it, and recommend ranked an older text-to-music member for "extend a song" at authoring time: an extension is not a txt2audio need, so do not take that pick.

Common mistakes

  • Hardcoding generative model IDs: availability differs per team and evolves. Re-discover each session, recommend for the need or search for a name; only the fixed first-party tool ids above stay constant.
  • Skipping model_schema_get: one audio model's parameters will not fit another (voices, durations, and lyric fields all differ).
  • Polling job_get in a loop: music jobs can run minutes. Use jobs_wait; on timeout re-call with pending_job_ids.
  • Pasting raw asset URLs into chat: use asset_display to play audio.
  • Passing format to asset_download for audio: it converts image formats only, so omit it.
  • Putting voice direction inside TTS text ("say this angrily"): direction can end up spoken. Use the schema's emotion or voice fields.
  • Writing a dialogue as one voice reading both parts: on a multi-speaker member, fill the speaker rows and prefix every turn with its label, since with the rows empty the member reads everything in its single voice.
  • Pasting lyrics into the style field: the model then describes a song instead of singing one.
  • Answering "make it longer" with a repaint or a new text-to-music run: the first keeps the duration, the second loses the song; the extend lane above keeps the slices the user chose.
  • Putting dry_run or wait inside parameters: they are model_run's own arguments, so a stray dry_run still charges and a stray wait blocks up to 180s.

© scenario-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/scenario-audio of scenario-labs/skills.

Open the folder on GitHubat commit f6f8ab7

Compare with similar skills

Scenario Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scenario Audio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scenario Audio this skillscenario-labs/skills946—~3kAutomated safety check: PassMIT
ShowtimeFavioVazquez/showtime220—~3kAutomated safety check: PassMIT
BlockrunBlockRunAI/blockrun-mcp391—~2.7kAutomated safety check: PassMIT
Audio And Videoglifxyz/glif-mcp-server213—~1.6kAutomated safety check: PassMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Showtime

    FavioVazquez/showtime

    A skill your agent uses when the user wants a video made, edited or finished: a launch or promo, product demo, explainer, trailer or teaser, tutorial or walkthrough, a screen recording turned into a…

    220 GitHub stars~3k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Blockrun

    BlockRunAI/blockrun-mcp

    Pay-per-call access to AI models, real-time data, media generation and multi-chain RPC over x402 micropayments (USDC on Base or Solana), or a BlockRun account API key.

    391 GitHub stars~2.7k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Audio And Video

    glifxyz/glif-mcp-server

    Make or edit audio and video with Glif, from a text brief or from a reference image, video or audio file.

    213 GitHub stars~1.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • Text To Sfx

    sonilo-ai/skills

    Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match.

    115 GitHub starsUsed in 1 repo~1.6k tokens
    Media & CreativeAuto-check: notes

More from scenario-labs/skills

All 146 skills in this repo
  • Scenario Blender Grease Pencil

    scenario-labs/skills

    A skill your agent uses when drawing or animating with Grease Pencil in Blender 5.x from Python: 2D or 2.5D illustration, frame-by-frame animation, a cutout or part-based 2D character, strokes with…

    946 GitHub stars~4.5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Blender Hair

    scenario-labs/skills

    A skill your agent uses when grooming hair or fur in Blender with hair curves, such as a character hairstyle, animal fur, procedural fur in geometry nodes, or hair cards and mesh hair for games.

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • A skill your agent uses when lighting, rendering or compositing in Blender: light a character, product or hero shot, interior at dusk or night, three-point or motivated lighting, sun and sky, HDRI…

    946 GitHub stars~5k tokensUpdated yesterday
    Auto-check passed
  • Scenario Chatgpt Pet Create

    scenario-labs/skills

    A skill your agent uses when creating a ChatGPT pet or Codex pet with Scenario: hatching an animated companion from a text idea, a character, mascot or brand cue, or reference photos and art; making…

    946 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Animation

    scenario-labs/skills

    A skill your agent uses when animating characters or scenes in Godot 4.7: AnimationPlayer clips and RESET, AnimationTree state machines and blend spaces built in code, Mixamo or glTF import, loop…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed
  • Scenario Godot Audio

    scenario-labs/skills

    A skill your agent uses when adding or fixing sound in Godot 4.7: audio buses and effects, volume sliders, 'too many sounds', combat audio with hundreds of enemies, sounds clipping or distorting, 3D…

    946 GitHub stars~4.7k tokensUpdated yesterday
    Auto-check passed

Questions about Scenario Audio

What does Scenario Audio do?

A skill your agent uses when generating or handling audio on Scenario via MCP. Scenario Audio is an agent skill from scenario-labs/skills. Use when generating or handling audio on Scenario via MCP.

When should I use Scenario Audio?

Scenario Audio fits situations like: handling audio on Scenario via MCP; include music tracks; full-length songs with vocals written from lyrics; background scores.

How do I install Scenario Audio in Claude Code?

Run `npx skills add scenario-labs/skills --skill scenario-audio -a claude-code`. Or copy the skill folder (skills/scenario-audio in scenario-labs/skills) into .claude/skills/scenario-audio in your project. Claude Code loads it when a task matches its description.

How do I install Scenario Audio in Codex?

Run `npx skills add scenario-labs/skills --skill scenario-audio -a codex`. Or copy the skill folder (skills/scenario-audio in scenario-labs/skills) into .agents/skills/scenario-audio in your project. Codex loads it when a task matches its description.

Can I use Scenario Audio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add scenario-labs/skills --skill scenario-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scenario-audio, .gemini/skills/scenario-audio, .github/skills/scenario-audio and .opencode/skills/scenario-audio in your project.

What does Scenario Audio need to run?

Going by SKILL.md and its folder, Scenario Audio needs the command-line tools its instructions call (curl and npx). Our summary lists: Node.js.

Does Scenario Audio access the network?

SKILL.md contains no URLs. Its commands use curl and npx, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Scenario Audio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scenario Audio use?

Scenario Audio is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scenario Audio use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scenario Audio?

Skills that share tags, products or a category with Scenario Audio: Showtime (FavioVazquez/showtime, 220 stars), Blockrun (BlockRunAI/blockrun-mcp, 391 stars), Audio And Video (glifxyz/glif-mcp-server, 213 stars) and Music (tadaspetra/loop, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scenario Audio?

scenario-labs (a GitHub organization) maintains it in scenario-labs/skills, which has 946 GitHub stars. The repository holds 146 skills in this directory. The repository was last updated on October 10, 2026.

Source: scenario-labs/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.