Stable Audio Prompting
nodetool-ai/nodetool
Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable…
A skill your agent uses for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds…
$ npx skills add calesthio/generative-media-skills --skill stable-audio -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install calesthio/generative-media-skills stable-audio --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/providers/sound-generation/stable-audio .claude/skills/stable-audio && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "stable-audio" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audio into .claude/skills/stable-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-audio", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audioType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add calesthio/generative-media-skills --skill stable-audio -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install calesthio/generative-media-skills stable-audio --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/providers/sound-generation/stable-audio .agents/skills/stable-audio && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "stable-audio" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audio into .agents/skills/stable-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-audio", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill stable-audio -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install calesthio/generative-media-skills stable-audio --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/providers/sound-generation/stable-audio .cursor/skills/stable-audio && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "stable-audio" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audio into .cursor/skills/stable-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-audio", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/calesthio/generative-media-skills.git --path skills/providers/sound-generation/stable-audio--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add calesthio/generative-media-skills --skill stable-audio -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install calesthio/generative-media-skills stable-audio --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/providers/sound-generation/stable-audio .gemini/skills/stable-audio && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "stable-audio" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audio into .gemini/skills/stable-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-audio", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install calesthio/generative-media-skills stable-audioInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add calesthio/generative-media-skills --skill stable-audio -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/providers/sound-generation/stable-audio .github/skills/stable-audio && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "stable-audio" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audio into .github/skills/stable-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-audio", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add calesthio/generative-media-skills --skill stable-audio -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install calesthio/generative-media-skills stable-audio --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/calesthio/generative-media-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/providers/sound-generation/stable-audio .opencode/skills/stable-audio && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "stable-audio" agent skill from https://github.com/calesthio/generative-media-skills/tree/main/skills/providers/sound-generation/stable-audio into .opencode/skills/stable-audio/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "stable-audio", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
stable-audioA skill your agent uses for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds…
Stable Audio is an agent skill from calesthio/generative-media-skills. Use for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds, stingers, and sonic-branding audio from text or source audio, planning rights-safe uploads, setting model parameters, polling asynchronous jobs, and reviewing generated audio for media projects.
Its SKILL.md is about 4.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `EVAL.md`).
It sits in Media & Creative, covering Music and audio generation and Async programming. It works with Stable Diffusion. The repository describes itself as: Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 8c85352. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
stability.aiarxiv.orgapi.stability.aigithub.comhuggingface.coplatform.stability.aiFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Stable Audio loads about 4.8k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 2,275 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from calesthio/generative-media-skills at commit 8c85352, republished under its MIT licence (© calesthio). 2,275 words, ~4,816 tokens.
.claude/skills/stable-audio/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Stable Audio is Stability AI's generative-audio family for music and sound-design outputs, not speech, voice cloning, dialogue, or final mix/mastering. Use it when a project needs instrumental music, loops, stems/ideas, sound effects, foley-like assets, ambience, stingers, transitions, or style/continuation edits from rights-cleared audio. Route speech, dubbing, voiceover, transcript alignment, and loudness mastering to the appropriate speech/audio-post tools instead.
All volatile facts in this guide were verified on 2026-07-10 from Stability AI official pages, the Stability API OpenAPI specification, official model cards/repositories, and Stability policy pages.
Documented facts:
/v2beta/audio/stable-audio-2/..., with model choices stable-audio-2 and stable-audio-2.5./v2beta/audio/stable-audio/..., with model=stable-audio-3.mp3 or wav output format.GET /v2beta/audio/results/{id} until HTTP 200 returns the audio.mp3 or wav, and requests are capped at 50 MB.mp3 or wav, and requests are capped at 100 MB.seed accepts 0 or omitted for random generation; explicit seeds may be recorded for reproducibility.steps differs by model: Stable Audio 2 accepts 30-100 steps and defaults to 50; Stable Audio 2.5 and Stable Audio 3 accept 4-8 steps and default to 8.cfg_scale ranges 1-25. API docs describe it as prompt-adherence strength; defaults are 7 for Stable Audio 2, 1 for Stable Audio 2.5, and 1 for Stable Audio 3. Treat it as a secondary control after prompt, duration, model, steps, and source-audio strategy.strength in audio-to-audio ranges 0-1 and controls source-audio influence. Stability describes 0 as identical to the input and 1 as equivalent to no audio influence.mask_start and mask_end in seconds to choose the replacement/continuation segment.credits = 17 + 0.06 * steps per successful generation, usually 20 credits at the default 50 steps and 23 credits at 100 steps; Stable Audio 2.5 is 20 credits per successful result; Stable Audio 3.0 is 26 credits per successful result. Stability states failed generations are not charged. Stability pricing uses 1 credit = $0.01 and is subject to change.small-music, small-sfx, and medium; the official repository lists Small music/SFX as CPU-capable 433M-parameter models with 120s max length and Medium as a 1.4B CUDA model with 380s max length. Stable Audio 3 Large is API-only in the official repo.Sources: Stability API reference / OpenAPI spec, Stable Audio 3 page, Stable Audio 2.0 announcement, Stable Audio 2.5 announcement, Stable Audio 3 repository, Stable Audio Open 1.0 model card, Stable Audio Open paper, Stable Audio 3 paper, Stability pricing page.
Use hosted Stable Audio 3 when the project needs current highest-capability hosted Stable Audio, up to ~6 minute beds, async generation is acceptable, or inpainting/continuation is part of the plan. It is the cleanest choice for longer music beds, adaptive game/music cues, sonic-branding variations, and complex ambience where 190 seconds is not enough.
Use hosted Stable Audio 2.5 when you need synchronous API behavior, 190 seconds is enough, and the brief is commercial music/sound production where Stable Audio 2.5's documented improvements in speed, musical structure, mood/genre prompt adherence, and inpainting matter.
Use hosted Stable Audio 2 only when a pipeline already depends on its behavior or when you intentionally need its 30-100 step range and cfg_scale behavior. Do not default to it just because the endpoint default is stable-audio-2; explicitly set the model.
Use Stable Audio 3 open weights when local or self-hosted generation, fine-tuning/LoRA experimentation, data-control review, or offline iteration matters. Choose:
small-music for quick CPU music drafts up to 120s;small-sfx for quick CPU sound-effect drafts up to 120s;medium for higher-quality local generation up to 380s when CUDA and dependency constraints are acceptable.Use Stable Audio Open 1.0 only when the 47s limit is acceptable and the workflow needs the older open model, diffusers/stable-audio-tools compatibility, or research comparison.
Do not use Stable Audio for:
Before any audio-to-audio or inpaint request, require the user or project brief to establish that every uploaded sound is rights-cleared for transformation. Stability's API docs and Stable Audio announcements state that copyrighted content is not allowed to be uploaded and describe content recognition/compliance scanning for uploads. Treat "I found this song online" as blocked until replaced with licensed, public-domain, commissioned, self-recorded, or otherwise authorized material.
For commercial media, record:
cfg_scale if used, strength, mask times, and output format;Stability's Terms of Service state that users are responsible for inputs and must have rights, licenses, and permissions for inputs; as between the user and Stability, Stability assigns any right it has in outputs subject to compliance and applicable law. The same terms also say outputs may be similar across users, users must verify legality/appropriateness before use, and Stability may use inputs/outputs to improve services unless the user opts out where available. Privacy/security policy claims should not be overstated: Stability states it uses organizational and technical safeguards, but no internet transmission/storage is guaranteed to be fully secure.
For confidential brand libraries, unreleased music, celebrity voices, minors, medical/legal content, or contractual IP, prefer enterprise-approved settings or local/open-weight routes only after the project's data-handling requirements are known. Do not upload sensitive stems just because the API is convenient.
Sources: Stability Terms of Service, Stability Acceptable Use Policy, Stability Privacy Policy, Stable Audio 2.0 announcement, Stable Audio 2.5 announcement, Stability SOC 2/SOC 3 announcement.
Stable Audio responds best when the prompt describes what the audio should sound like, not what the video should show. Translate visual intent into sonic parameters.
Documented prompt elements from Stability guidance include genre/subgenre, style, tempo/BPM, mood, and instrument type. Production heuristics below are not documented guarantees; use them because they make review and iteration easier:
Keep prompts internally consistent. A single prompt that asks for "minimal ambient corporate piano" and "aggressive drum & bass festival drop" will produce less controllable review targets.
Duration:
Model:
model explicitly. Do not rely on endpoint defaults.stable-audio-3 for 191-380s hosted work or async pipelines.stable-audio-2.5 for synchronous 1-190s hosted work unless project history requires 2.0.Steps:
Seed:
Output format:
wav for assets that will be edited, looped, layered, mixed, or mastered.mp3 for quick previews, low-risk temp tracks, or delivery systems that require it.Audio-to-audio strength:
Inpaint masks:
mask_start slightly before the flawed region and mask_end slightly after it so the model has room to blend.Text-to-audio music bed:
wav with documented seed/parameters.Text-to-audio SFX/foley:
Audio-to-audio transformation:
Inpainting/continuation:
Intent: an optimistic bed under narration for a SaaS launch reel, no vocals, easy to duck.
Route: hosted Stable Audio 2.5 text-to-audio, synchronous, wav.
Parameters:
endpoint: POST /v2beta/audio/stable-audio-2/text-to-audio
model: stable-audio-2.5
duration: 40
steps: 8
output_format: wav
seed: 0Prompt:
40-second modern product-launch instrumental bed at 104 BPM, optimistic and confident, clean electronic pop with warm analog synth pulses, soft piano accents, light brushed percussion, subtle bass, wide polished stereo mix. Structure: 4-second gentle intro, steady lift through 25 seconds, restrained final button ending. Designed to sit under spoken narration. No vocals, no lead guitar solo, no aggressive drums, no recognizable melody.Why structured this way: The prompt names function, duration, BPM, mood, arrangement, section behavior, mix priority, and exclusions relevant to narration.
Expected review: Check the first 4 seconds for usable intro space, verify the 25-30s region can support the hero claim, and cut/fade the 40s render to the final 30s edit.
Likely failures: too much lead melody, drums masking speech, unresolved ending, or tempo not matching the edit. Iterate by reducing instrumentation or asking for "more sparse, more sidechain space for narration."
Intent: transform a client-owned two-note chime into a family of softer onboarding sounds.
Precondition: project log confirms the client owns the chime and permits transformation.
Route: hosted Stable Audio 3 audio-to-audio, async, wav.
Parameters:
endpoint: POST /v2beta/audio/stable-audio/audio-to-audio
model: stable-audio-3
duration: 8
steps: 8
strength: 0.42
output_format: wav
seed: 0Prompt:
8-second soft premium onboarding chime derived from the source gesture, calm and reassuring, two-note identity preserved as a subtle motif, warm glass mallet tone layered with quiet felt piano resonance, gentle airy tail, close clean studio sound, no melody beyond the two-note motif, no voice, no percussion, no harsh digital sparkle.Workflow: Submit the job, store the returned generation id, poll /v2beta/audio/results/{id}, save the returned seed and request id, then create a short candidate sheet with waveform, loudness, and subjective notes.
Likely failures: source motif disappears at high strength; source is too literal at low strength; tail is too long for UI. Adjust strength first, not prompt complexity.
Intent: replace a 12-second section with accidental foreground chatter in an otherwise useful 90-second sci-fi hallway ambience.
Route: hosted Stable Audio 2.5 inpaint, synchronous, wav.
Parameters:
endpoint: POST /v2beta/audio/stable-audio-2/inpaint
model: stable-audio-2.5
duration: 90
steps: 8
mask_start: 31.5
mask_end: 45.0
output_format: wav
seed: keep original asset seed if known; otherwise 0Prompt:
90-second seamless sci-fi hallway ambience, low ventilation rumble, distant electrical hum, subtle metallic room tone, occasional very soft servo movement far away, tense but not musical. The replacement section should match the surrounding ambience and contain no speech, no footsteps, no alarms, no rhythmic music.Review: Listen across 28-48s on headphones for seam clicks, room-tone jump, new foreground events, and stereo image shifts. If the new section draws attention, reduce event density in the prompt and widen the mask slightly.
Before accepting a Stable Audio asset:
wav for production; derive compressed preview/delivery formats from the approved master.© calesthio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/providers/sound-generation/stable-audio of calesthio/generative-media-skills.
Open the folder on GitHubat commit 8c85352
Stable Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Stable Audio this skillcalesthio/generative-media-skills | 193 | — | ~4.8k | Automated safety check: Pass | MIT | |
| Stable Audio Promptingnodetool-ai/nodetool | 560 | — | ~1.8k | Automated safety check: Pass | AGPL-3.0 | |
| Routerbase Media Generationaiskillstore/marketplace | 430 | — | ~972 | Automated safety check: Pass | None | |
| Musictadaspetra/loop | 296 | 2 repos | ~827 | Automated safety check: Pass | MIT | |
| Sound Effectstadaspetra/loop | 296 | 2 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Text To Sfxsonilo-ai/skills | 115 | 1 repos | ~1.6k | Automated safety check: Notes | MIT |
nodetool-ai/nodetool
Prompt Stability's Stable Audio line — the genre/instruments/mood/BPM order its training metadata expects, the TrackType and VocalType tags that separate music, stems and sound effects on Stable…
aiskillstore/marketplace
Build image, video, and audio generation workflows on RouterBase.
tadaspetra/loop
Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.
tadaspetra/loop
Generate sound effects from text descriptions using ElevenLabs.
sonilo-ai/skills
Generate a sound effect from a text description using Sonilo — a UI chime, a whoosh, an impact, ambience, a stylized cue — when there is no video to match.
T8mars/T8-penguin-canvas
Turn a brief music description and optional tagged lyrics into a professional MiniMax Music 3 structured caption with Global Metadata, Vocal Details, and a section-aware Arrangement.
calesthio/generative-media-skills
A skill your agent uses to turn generated, captured, scanned, or modeled 3D output into production-ready standalone assets for DCC, real-time engine, web, or interchange delivery.
calesthio/generative-media-skills
Provider-independent audio mixing and mastering direction for AI agents finishing generated videos, ads, trailers, explainers, podcasts, recuts, avatar clips, music videos, documentaries, and social…
calesthio/generative-media-skills
Provider-independent captions and media accessibility direction for AI agents producing or finishing generated videos, ads, social clips, explainers, avatar videos, documentaries, podcasts/video…
calesthio/generative-media-skills
Provider-independent production workflow for agents assembling, auditing, executing, and handing off ComfyUI node-graph workflows for image, video, upscale, inpaint, conditioning, and batch media…
calesthio/generative-media-skills
Provider-independent FFmpeg finishing workflow for AI agents preparing generated or edited media deliverables.
calesthio/generative-media-skills
Provider-independent quality assurance for AI-generated and AI-assisted media.
Works with
Categories
A skill your agent uses for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds…. Stable Audio is an agent skill from calesthio/generative-media-skills. Use for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds, stingers, and sonic-branding audio from text or source audio, planning rights-safe uploads, setting model parameters, polling asynchronous jobs, and reviewing generated audio for media projects.
Stable Audio fits situations like: stability AI Stable Audio production work: selecting Stable Audio hosted API; open-weight models; sonic-branding audio from text; planning rights-safe uploads.
Run `npx skills add calesthio/generative-media-skills --skill stable-audio -a claude-code`. Or copy the skill folder (skills/providers/sound-generation/stable-audio in calesthio/generative-media-skills) into .claude/skills/stable-audio in your project. Claude Code loads it when a task matches its description.
Run `npx skills add calesthio/generative-media-skills --skill stable-audio -a codex`. Or copy the skill folder (skills/providers/sound-generation/stable-audio in calesthio/generative-media-skills) into .agents/skills/stable-audio in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/generative-media-skills --skill stable-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/stable-audio, .gemini/skills/stable-audio, .github/skills/stable-audio and .opencode/skills/stable-audio in your project.
SKILL.md names no scripts, command-line tools or credentials: Stable Audio is instructions for the agent only.
SKILL.md names 6 domains. As links in the text: stability.ai, arxiv.org, api.stability.ai, github.com, huggingface.co and platform.stability.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Stable Audio is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.8k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Stable Audio: Stable Audio Prompting (nodetool-ai/nodetool, 560 stars), Routerbase Media Generation (aiskillstore/marketplace, 430 stars), Music (tadaspetra/loop, 296 stars) and Sound Effects (tadaspetra/loop, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
calesthio (a GitHub user) maintains it in calesthio/generative-media-skills, which has 193 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on July 14, 2026.
Source: calesthio/generative-media-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.