Agent skill

Stable Audio Prompting

by nodetool-ai in nodetool-ai/nodetool

Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable…

AGPL-3.0Auto-check passedMedia & Creative

Install Stable Audio Prompting

skills CLI
$ npx skills add nodetool-ai/nodetool --skill stable-audio-prompting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nodetool-ai/nodetool stable-audio-prompting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/system-skills/stable-audio-prompting .claude/skills/stable-audio-prompting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
stable-audio-prompting
GitHub stars
560
Token cost
~1.8k tokens
SKILL.md length
928 words
Files
1
Skills in repo
127
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable…

  • Works in 4 steps: Genre — the core style, first. → Instruments — with character, not just… → Mood and production — the feel, plus the… → …
  • The model id names Stable Audio (fal-ai/stable-audio-25/text-to-audio
  • SKILL.md covers Music prompts, Stems and single instruments, Sound effects and Duration and parameters, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Stable Audio Prompting is an agent skill from nodetool-ai/nodetool. Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable Audio 3, matching duration to what you described, which checkpoints guidancescale and step count actually affect, and prompting the inpaint, outpaint and audio-to-audio variants where the surrounding audio outranks the prompt. Use whenever the model id names Stable Audio (fal-ai/stable-audio-25/text-to-audio…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Music and audio generation. It works with fal and Stable Diffusion. The repository describes itself as: Agent-first Creative Workspace. The licence is AGPL-3.0.

When your agent uses it

  • The model id names Stable Audio (fal-ai/stable-audio-25/text-to-audio
  • Fal-ai/stable-audio-3/medium/text-to-audio
  • Fal-ai/stable-audio-3/small/music/text-to-audio
  • Fal-ai/stable-audio-3/small/sfx/text-to-audio

Example prompts

  • “/stable-audio-prompting”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Genre — the core style, first.
  2. Instruments — with character, not just names. "Smooth electric piano",
  3. Mood and production — the feel, plus the production era or treatment.
  4. BPM — and the key when it matters.

What it can do on your machine

Read from SKILL.md and the folder at commit 339f069. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • stability.ai
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Stable Audio Prompting loads about 1.8k tokens when it runs. Until then it costs about 206 tokens; SKILL.md has 928 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~206
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nodetool-ai/nodetool at commit 339f069, republished under its AGPL-3.0 licence (© nodetool-ai). 928 words, ~1,765 tokens.

Download SKILL.mdSave it as .claude/skills/stable-audio-prompting/SKILL.md (or your agent's skills folder).
name
stable-audio-prompting
description
Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable Audio 3, matching duration to what you described, which checkpoints guidance_scale and step count actually affect, and prompting the inpaint, outpaint and audio-to-audio variants where the surrounding audio outranks the prompt. Use whenever the model id names Stable Audio (fal-ai/stable-audio-25/text-to-audio, fal-ai/stable-audio-3/medium/text-to-audio, fal-ai/stable-audio-3/small/music/text-to-audio, fal-ai/stable-audio-3/small/sfx/text-to-audio, fal-ai/stable-audio-3/medium/audio-to-audio, fal-ai/stable-audio-25/inpaint, stability-ai/stable-audio-2.5) on generate_music or an audio generation node.

Stable Audio → write the metadata, not a mood

Stable Audio was trained on library-music metadata, so it responds to prompts shaped like library-music metadata: what the genre is, what is playing, how it feels, how fast it goes. Poetic descriptions of a feeling get a generic bed.

Reach the music endpoints with find_model for text_to_music, then generate_music. The SFX checkpoints and the audio-to-audio, inpaint and outpaint variants are not in the music catalog — run those nodes with search_nodes and invoke_node.

Music prompts

Four elements, in this order:

  1. Genre — the core style, first.
  2. Instruments — with character, not just names. "Smooth electric piano", "textural percussion", "distorted bass" carry more than "piano, drums, bass".
  3. Mood and production — the feel, plus the production era or treatment.
  4. BPM — and the key when it matters.

An exciting breakbeat instrumental for a fast-paced game, funky electric guitar chords, steady break drums, smooth electric piano and supporting bass. Fresh, modern and adventurous, 105 BPM.

Two habits that measurably help: vary the vocabulary instead of repeating the same adjective, and reference an era to describe production — "80s gated reverb", "90s grunge distortion" — which encodes a whole chain in three words.

On Stable Audio 3, prefix the type. TrackType: Music, VocalType: Instrumental is the tag pair for instrumental beds and is worth setting explicitly; the model otherwise decides whether to bring in a vocal.

Stems and single instruments

Start with TrackType: Instrument, then the instrument, genre, mood and BPM. This is the path for something that has to sit in a mix rather than be the mix. Playing technique, recording environment and effects all read: "close-mic'd upright bass, fingered, small wooden room, light tape saturation".

Sound effects

TrackType: SFX, then three things:

  1. Source — the object or instrument making the sound.
  2. Action — how it is triggered, how long it rings, how it decays.
  3. Production — mic placement, room character, processing.

TrackType: SFX. A steel toolbox lid slamming shut on a concrete workshop floor, sharp metallic impact with a short rattling decay, close mic, small room with a hard early reflection.

Then set a short duration. Most effects are under two seconds, and a 30-second request for a door slam gets you a door slam followed by 28 seconds of room.

Duration and parameters

Match the duration to what you described. The Stable Audio 3 medium checkpoint goes to 380 seconds, but a prompt describing a loop or a hit produces its best result when the length fits the description; for a music bed, 60–90 seconds is the range the guide recommends. Note the field changes name — seconds_total on 2.5, duration on 3 — and 2.5 defaults to 190 seconds, which is three minutes of audio nobody asked for.

ParameterWhat it doesWhere to sit
num_inference_stepsSampling stepsThe distilled checkpoints are tuned for the default 8 and gain little above it; raise it only on a base checkpoint
guidance_scalePrompt adherenceOn Stable Audio 3, effective only on base (non-distilled) checkpoints — raising it on a distilled one changes nothing
negative_promptQualities to avoid (Stable Audio 3 only)Name the artifact you hear: "clipping", "muddy low end", "vocal chops"
enable_prompt_expansionLLM prompt rewriteOff by default; useful for a three-word prompt, harmful once your prompt is already metadata
seedComparabilityPin it while you change one element at a time

The base versus distilled split is the one that catches people: on a distilled Stable Audio 3 checkpoint the two dials most people reach for do nothing, and the prompt is the whole instrument.

Show full SKILL.md (340 more words)Show less

Editing existing audio

Audio-to-audio seeds generation from a clip, and the dial has two names. Stable Audio 3 takes init_noise_level (0.9 default): 0.1 keeps the source close, 1.0 replaces it with noise and generates outright. Stable Audio 2.5 takes strength (0.8 default), which runs the same way — 0 returns the input. Low keeps melody and rhythm; high strips them and keeps only broad character.

Inpaint and outpaint hold everything outside a masked region and regenerate inside it — mask_start_seconds / mask_end_seconds on Stable Audio 3, mask_start / mask_end on 2.5. Here the surrounding audio matters more than the prompt: a short mask is pulled hard toward its context regardless of what you wrote, and only a wide mask gives the prompt room. If an inpaint is ignoring you, widen the mask before rewriting the prompt.

Symptoms

What went wrongWhat to change
A generic bedRewrite as genre, instruments with character, mood, BPM
An unwanted vocalAdd TrackType: Music, VocalType: Instrumental to the prompt
The effect is padded with room toneCut the duration to the length of the actual event
Raising guidance changed nothingYou are on a distilled checkpoint — switch to base or fix the prompt
An inpaint ignores the promptWiden the masked region
Two takes are not comparablePin the seed and change one element

Check the result with analyze_audio, analyze_audio_spectrum and detect_audio_events rather than listening for it — a missing instrument or a clipped peak shows up there faster than by ear.

Where it lands

A bed from here is the "track of your own" build in video-audio-continuity: one clip on its own audio track for the whole runtime, with the generated shots muted under it. Because the prompt states the BPM, beat-sync-editing starts with a grid it can predict and then confirms with detect_audio_events on the file. A TrackType: SFX hit of 1–3 s is what logo-reveal lands a mark on. motion-graphics carries the timeline ops that lay the clip down.

Adapted from Stability's Stable Audio 2.5 prompt guide and the Stable Audio 3 prompting guide: https://stability.ai/implementations/stable-audio-25-prompt-guide https://github.com/Stability-AI/stable-audio-3/blob/main/docs/guides/prompting.md

© nodetool-ai, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/system-skills/stable-audio-prompting of nodetool-ai/nodetool.

Open the folder on GitHubat commit 339f069

Compare with similar skills

Stable Audio Prompting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Stable Audio Prompting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Stable Audio Prompting this skillnodetool-ai/nodetool560—~1.8kAutomated safety check: PassAGPL-3.0
Keirouter Imagemydisha/keirouter147—~691Automated safety check: PassMIT
Fal Generatenexu-io/open-design100k—~306Automated safety check: PassApache-2.0
Videoagent Audio Studiopexoai/pexo-skills804—~1.7kAutomated safety check: PassMIT
Fal AIhoodini/ai-agents-skills282—~2.1kAutomated safety check: NotesNone
Moonvalley Mareycalesthio/generative-media-skills197—~5.2kAutomated safety check: PassMIT

Similar skills

  • Keirouter Image

    mydisha/keirouter

    Generate images via KeiRouter /v1/images/generations using OpenAI DALL-E / Gemini Imagen / FLUX / MiniMax / Stability AI / Fal.ai models.

    147 GitHub stars~691 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Fal Generate

    nexu-io/open-design

    Generate images and videos using fal.ai AI models. An agent skill from nexu-io/open-design.

    100k GitHub stars~306 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Videoagent Audio Studio

    pexoai/pexo-skills

    Tired of juggling multiple audio APIs?. An agent skill from pexoai/pexo-skills.

    804 GitHub stars~1.7k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Fal AI

    hoodini/ai-agents-skills

    Generate images, videos, and audio with fal.ai serverless AI.

    282 GitHub stars~2.1k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes
  • Moonvalley Marey

    calesthio/generative-media-skills

    Produce video with Moonvalley's Marey model family (Marey Realism v1.5) — a filmmaker-oriented, 1080p/24fps generative video model marketed as trained exclusively on licensed data.

    197 GitHub stars~5.2k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Stable Audio

    calesthio/generative-media-skills

    A skill your agent uses for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds…

    197 GitHub stars~4.8k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

More from nodetool-ai/nodetool

All 127 skills in this repo
  • Beat Sync Editing

    nodetool-ai/nodetool

    Cut a NodeTool timeline to music and shape its pacing — detect the beat grid, place cuts on phrases, pick a cut type, build speed ramps with time remap, and give the piece an arc.

    560 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Caption Titles

    nodetool-ai/nodetool

    Add and animate a consistent text layer on an existing NodeTool timeline.

    560 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Color Motion

    nodetool-ai/nodetool

    Choose and animate colour on a NodeTool timeline, including shape and text gradients, colour grades, 3D LUTs, and dither.

    560 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Commercial Beat Sheet

    nodetool-ai/nodetool

    Write a shootable, precisely timed commercial beat sheet and store it as a NodeTool storyboard, with a consistent entity roster behind every shot.

    560 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Elevenlabs Audio Prompting

    nodetool-ai/nodetool

    Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…

    560 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Frame Composition

    nodetool-ai/nodetool

    Stage the frame on a NodeTool timeline — grids, focal placement, safe areas per aspect ratio, depth layers and parallax, camera moves, and where elements enter and leave.

    560 GitHub stars~3.5k tokensUpdated today
    Auto-check passed

Questions about Stable Audio Prompting

What does Stable Audio Prompting do?

Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable…. Stable Audio Prompting is an agent skill from nodetool-ai/nodetool. Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable Audio 3, matching duration to what you described, which checkpoints guidancescale and step count actually affect, and prompting the inpaint, outpaint and audio-to-audio variants where the surrounding audio outranks the prompt.

When should I use Stable Audio Prompting?

Stable Audio Prompting fits situations like: the model id names Stable Audio (fal-ai/stable-audio-25/text-to-audio; fal-ai/stable-audio-3/medium/text-to-audio; fal-ai/stable-audio-3/small/music/text-to-audio; fal-ai/stable-audio-3/small/sfx/text-to-audio.

How do I install Stable Audio Prompting in Claude Code?

Run `npx skills add nodetool-ai/nodetool --skill stable-audio-prompting -a claude-code`. Or copy the skill folder (packages/system-skills/stable-audio-prompting in nodetool-ai/nodetool) into .claude/skills/stable-audio-prompting in your project. Claude Code loads it when a task matches its description.

How do I install Stable Audio Prompting in Codex?

Run `npx skills add nodetool-ai/nodetool --skill stable-audio-prompting -a codex`. Or copy the skill folder (packages/system-skills/stable-audio-prompting in nodetool-ai/nodetool) into .agents/skills/stable-audio-prompting in your project. Codex loads it when a task matches its description.

Can I use Stable Audio Prompting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nodetool-ai/nodetool --skill stable-audio-prompting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stable-audio-prompting, .gemini/skills/stable-audio-prompting, .github/skills/stable-audio-prompting and .opencode/skills/stable-audio-prompting in your project.

What does Stable Audio Prompting need to run?

SKILL.md names no scripts, command-line tools or credentials: Stable Audio Prompting is instructions for the agent only.

Does Stable Audio Prompting access the network?

SKILL.md names 2 domains. As links in the text: stability.ai and github.com. This is read from the text; nothing was executed.

Is Stable Audio Prompting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Stable Audio Prompting use?

Stable Audio Prompting is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Stable Audio Prompting use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Stable Audio Prompting?

Skills that share tags, products or a category with Stable Audio Prompting: Keirouter Image (mydisha/keirouter, 147 stars), Fal Generate (nexu-io/open-design, 100k stars), Videoagent Audio Studio (pexoai/pexo-skills, 804 stars) and Fal AI (hoodini/ai-agents-skills, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Stable Audio Prompting?

nodetool-ai (a GitHub organization) maintains it in nodetool-ai/nodetool, which has 560 GitHub stars. The repository holds 127 skills in this directory. The repository was last updated on October 10, 2026.

Source: nodetool-ai/nodetool on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.