Agent skill

Elevenlabs Audio Prompting

by nodetool-ai in nodetool-ai/nodetool

Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…

AGPL-3.0Auto-check passedMedia & Creative

Install Elevenlabs Audio Prompting

skills CLI
$ npx skills add nodetool-ai/nodetool --skill elevenlabs-audio-prompting -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nodetool-ai/nodetool elevenlabs-audio-prompting --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nodetool-ai/nodetool.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/system-skills/elevenlabs-audio-prompting .claude/skills/elevenlabs-audio-prompting && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
elevenlabs-audio-prompting
GitHub stars
560
Token cost
~1.9k tokens
SKILL.md length
959 words
Files
1
Skills in repo
127
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…

  • The model id names an ElevenLabs generation endpoint (fal-ai/elevenlabs/tts/eleven-v3
  • SKILL.md covers Voice first, then tags, Stability is the delivery dial, Audio tags and Several speakers, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Fal-ai/elevenlabs/tts/multilingual-v2

What it does

Elevenlabs Audio Prompting is an agent skill from nodetool-ai/nodetool. Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead of SSML for pacing, the per-line dialogue format for several speakers, source-action-space sound-effect briefs with promptinfluence and looping, and the music composition plan that pins a song's sections. Use whenever the model id names an ElevenLabs generation endpoint (fal-ai/elevenlabs/tts/eleven-v3…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice and Music and audio generation. It works with ElevenLabs and fal. The repository describes itself as: Agent-first Creative Workspace. The licence is AGPL-3.0.

When your agent uses it

  • The model id names an ElevenLabs generation endpoint (fal-ai/elevenlabs/tts/eleven-v3
  • Fal-ai/elevenlabs/tts/multilingual-v2
  • Fal-ai/elevenlabs/text-to-dialogue/eleven-v3
  • Fal-ai/elevenlabs/sound-effects/v2

Example prompts

  • “/elevenlabs-audio-prompting”

What it can do on your machine

Read from SKILL.md and the folder at commit fefb6d1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • elevenlabs.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Elevenlabs Audio Prompting loads about 1.9k tokens when it runs. Until then it costs about 235 tokens; SKILL.md has 959 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~235
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nodetool-ai/nodetool at commit fefb6d1, republished under its AGPL-3.0 licence (© nodetool-ai). 959 words, ~1,927 tokens.

Download SKILL.mdSave it as .claude/skills/elevenlabs-audio-prompting/SKILL.md (or your agent's skills folder).
name
elevenlabs-audio-prompting
description
Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead of SSML for pacing, the per-line dialogue format for several speakers, source-action-space sound-effect briefs with prompt_influence and looping, and the music composition plan that pins a song's sections. Use whenever the model id names an ElevenLabs generation endpoint (fal-ai/elevenlabs/tts/eleven-v3, fal-ai/elevenlabs/tts/multilingual-v2, fal-ai/elevenlabs/text-to-dialogue/eleven-v3, fal-ai/elevenlabs/sound-effects/v2, fal-ai/elevenlabs/music, elevenlabs/text-to-speech-multilingual-v2, elevenlabs/text-to-dialogue-v3, elevenlabs/v2-multilingual, elevenlabs/turbo-v2.5) on generate_speech, generate_music or a TextToSpeech node. Not for the transcription, dubbing or isolation endpoints, which take no prompt.

ElevenLabs → direct the read, don't just submit the text

The text is the smallest part of an ElevenLabs request. What decides whether it sounds directed is the voice you picked, the stability you set, and the tags and punctuation carrying the performance.

Reach speech with find_model for text_to_speech, then generate_speech; music with find_model for text_to_music, then generate_music. The sound-effects endpoint is not in either catalog — run its node with search_nodes and invoke_node.

Voice first, then tags

A tag only does what the voice can already do. [shouts] on a voice trained on soft narration produces a slightly louder narration. Pick a voice whose range covers the delivery you want, then direct within it.

  • Emotionally varied voices take direction across a scene.
  • Narrow, consistent voices are right when every line is the same register.
  • Neutral voices are the most stable across languages and styles.

Stability is the delivery dial

The single most consequential setting. ElevenLabs exposes it as three named settings; the numeric APIs take the same three values, and the dialogue endpoint rounds anything else to the nearest of them.

ValueBehaviourUse for
0.0 — creativeMost expressive, most responsive to tags, occasional hallucinationCharacter work, dialogue, anything with [whispers] or [laughs]
0.5 — naturalBalanced, closest to the reference recordingNarration, most production reads
1.0 — robustVery consistent, largely ignores directional tagsLong documents, repeated renders that must match

At 1.0 your tags stop working. If a tag is being ignored, check stability before rewriting the tag.

Audio tags

Bracketed, inline, and they act from that point in the line. They are a v3 feature: on multilingual-v2, turbo and flash there is nothing to act on them, and punctuation plus sentence structure are the only direction you have.

  • Delivery: [whispers], [shouts], [sarcastic], [curious], [excited], [mischievously]
  • Reactions: [laughs], [sighs], [crying], [clears throat], [snorts]
  • Environment: [applause], [clapping], [gunshot], [explosion]
  • Experimental: [sings], [strong French accent]

Punctuation carries the rest, and v3 does not support SSML break tags. Ellipses are pauses and hesitation, capitals are emphasis, ordinary commas and full stops set the rhythm:

[sighs] It was a VERY long day … nobody listens any more. [quietly] Maybe that's the point.

Give the model a few sentences rather than a fragment. Short isolated lines deliver inconsistently because there is no context to read the register from. language_code forces a language when the text alone is ambiguous.

Several speakers

The dialogue endpoint takes a list of inputs, each with its own text and voice, rather than one block of text with names in it. Tag each line for its own delivery, and write interruptions as they happen:

[Ana, voice A]  "You said you'd call." [flat]
[Ruben, voice B] [defensive] "I did — twice —"
[Ana, voice A]  [cutting in] "Once. And you hung up."

Distinct voices per speaker is what makes it a conversation; the same voice twice reads as one person talking to themselves.

Sound effects

Brief them as source, action and space: what makes the sound, what happens to it, and the room it happens in.

A heavy oak door swinging shut and latching in a stone corridor, long natural reverb tail, close mic on the latch.

  • duration_seconds runs 0.5–22; leave it unset and the model infers a length from the prompt. Impacts want 1–3 s, beds want 10–20 s.
  • prompt_influence defaults to 0.3. Raise it toward 1 to follow the brief closely and lose variation; lower it when you want takes to choose from.
  • loop makes the tail blend into the head — the setting for ambience and game audio, and pointless for a one-shot.
Show full SKILL.md (395 more words)Show less

Music

A prompt gets you a track: genre, instrumentation with character, tempo in bpm, emotional arc, and what it is for.

Fast-paced electronic chase cue for a game trailer, driving synth arpeggios, punchy drums, distorted bass, rising tension with abrupt transitions, 130–150 bpm.

When the structure matters more than the vibe, send a composition_plan instead: an ordered list of sections, each with its own durationMs, positiveStyles and negativeStyles. That is what makes an intro stay sparse and a drop actually land where you wanted it. respect_sections_durations enforces those lengths; music_length_ms applies only to the prompt path. force_instrumental guarantees no vocal — without it, a prompt that does not mention vocals may still come back sung.

Naming an artist, a band or copyrighted lyrics is rejected outright; the error carries a rephrasing suggestion. Describe the sound instead of the reference.

Symptoms

What went wrongWhat to change
Tags do nothingLower stability to 0.5 or 0.0, and check the voice can do that delivery
The read is flatAdd punctuation and a delivery tag; stop relying on the words alone
Pauses are ignoredUse ellipses and line structure — SSML breaks are not supported
Two speakers sound the sameAssign a different voice per inputs entry
The effect is too variableRaise prompt_influence, and set an explicit duration
A loop clicksSet loop true and regenerate rather than trimming
A song ignores its structureMove from a prompt to a composition_plan
The prompt was refusedRemove artist and band names; describe the sound

Check the result with analyze_audio and detect_audio_events rather than listening through it, and transcribe_audio when you need to prove the words landed as written.

Where it lands

A voice or a bed from here is the "track of your own" build in video-audio-continuity: it goes on its own audio track and the generated clips are muted under it, which is what lets a multi-scene piece keep one continuous mix. For a storyboard's script, voice_script_lines voices every line in one call, each with its own voice, instead of one generate_speech per line. A bed built from a composition_plan has sections of known length, which is the grid beat-sync-editing cuts to; a sound-effect hit is what logo-reveal times a sting against. motion-graphics carries the timeline ops that lay the clip down.

Adapted from the ElevenLabs prompting documentation for Eleven v3, sound effects and music: https://elevenlabs.io/docs/best-practices/prompting/eleven-v3 https://elevenlabs.io/docs/overview/capabilities/sound-effects https://elevenlabs.io/docs/eleven-api/guides/cookbooks/music.md

© nodetool-ai, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/system-skills/elevenlabs-audio-prompting of nodetool-ai/nodetool.

Open the folder on GitHubat commit fefb6d1

Compare with similar skills

Elevenlabs Audio Prompting next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Elevenlabs Audio Prompting compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Elevenlabs Audio Prompting this skillnodetool-ai/nodetool560—~1.9kAutomated safety check: PassAGPL-3.0
Videoagent Audio Studiopexoai/pexo-skills804—~1.7kAutomated safety check: PassMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
ElevenLabs Voiceover Generatordigitalsamba/claude-code-video-toolkit2.2k1 repos~2.7kAutomated safety check: NotesMIT
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT
Hyperframes Mediachmonitor/chmonitor3011 repos~2.8kAutomated safety check: NotesGPL-3.0

Similar skills

  • Videoagent Audio Studio

    pexoai/pexo-skills

    Tired of juggling multiple audio APIs?. An agent skill from pexoai/pexo-skills.

    804 GitHub stars~1.7k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    301 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • Fal AI

    mikeOnBreeze/cc-crossbeam

    This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API.

    293 GitHub stars~1.9k tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes

More from nodetool-ai/nodetool

All 127 skills in this repo
  • Beat Sync Editing

    nodetool-ai/nodetool

    Cut a NodeTool timeline to music and shape its pacing — detect the beat grid, place cuts on phrases, pick a cut type, build speed ramps with time remap, and give the piece an arc.

    560 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Caption Titles

    nodetool-ai/nodetool

    Add and animate a consistent text layer on an existing NodeTool timeline.

    560 GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Color Motion

    nodetool-ai/nodetool

    Choose and animate colour on a NodeTool timeline, including shape and text gradients, colour grades, 3D LUTs, and dither.

    560 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Commercial Beat Sheet

    nodetool-ai/nodetool

    Write a shootable, precisely timed commercial beat sheet and store it as a NodeTool storyboard, with a consistent entity roster behind every shot.

    560 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Frame Composition

    nodetool-ai/nodetool

    Stage the frame on a NodeTool timeline — grids, focal placement, safe areas per aspect ratio, depth layers and parallax, camera moves, and where elements enter and leave.

    560 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Gpt Image 2 Prompting

    nodetool-ai/nodetool

    Prompt OpenAI's GPT Image 2 — the five-slot Scene/Subject/Details/Use case/Constraints template it responds to, the change-versus-preserve shape for edits, labelled multi-image compositing, and how…

    560 GitHub stars~1.6k tokensUpdated today
    Auto-check passed

Works with

Questions about Elevenlabs Audio Prompting

What does Elevenlabs Audio Prompting do?

Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…. Elevenlabs Audio Prompting is an agent skill from nodetool-ai/nodetool. Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead of SSML for pacing, the per-line dialogue format for several speakers, source-action-space sound-effect briefs with promptinfluence and looping, and the music composition plan that pins a song's sections.

When should I use Elevenlabs Audio Prompting?

Elevenlabs Audio Prompting fits situations like: the model id names an ElevenLabs generation endpoint (fal-ai/elevenlabs/tts/eleven-v3; fal-ai/elevenlabs/tts/multilingual-v2; fal-ai/elevenlabs/text-to-dialogue/eleven-v3; fal-ai/elevenlabs/sound-effects/v2.

How do I install Elevenlabs Audio Prompting in Claude Code?

Run `npx skills add nodetool-ai/nodetool --skill elevenlabs-audio-prompting -a claude-code`. Or copy the skill folder (packages/system-skills/elevenlabs-audio-prompting in nodetool-ai/nodetool) into .claude/skills/elevenlabs-audio-prompting in your project. Claude Code loads it when a task matches its description.

How do I install Elevenlabs Audio Prompting in Codex?

Run `npx skills add nodetool-ai/nodetool --skill elevenlabs-audio-prompting -a codex`. Or copy the skill folder (packages/system-skills/elevenlabs-audio-prompting in nodetool-ai/nodetool) into .agents/skills/elevenlabs-audio-prompting in your project. Codex loads it when a task matches its description.

Can I use Elevenlabs Audio Prompting in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nodetool-ai/nodetool --skill elevenlabs-audio-prompting -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/elevenlabs-audio-prompting, .gemini/skills/elevenlabs-audio-prompting, .github/skills/elevenlabs-audio-prompting and .opencode/skills/elevenlabs-audio-prompting in your project.

What does Elevenlabs Audio Prompting need to run?

SKILL.md names no scripts, command-line tools or credentials: Elevenlabs Audio Prompting is instructions for the agent only.

Does Elevenlabs Audio Prompting access the network?

SKILL.md names 1 domain. As links in the text: elevenlabs.io. This is read from the text; nothing was executed.

Is Elevenlabs Audio Prompting safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Elevenlabs Audio Prompting use?

Elevenlabs Audio Prompting is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Elevenlabs Audio Prompting use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Elevenlabs Audio Prompting?

Skills that share tags, products or a category with Elevenlabs Audio Prompting: Videoagent Audio Studio (pexoai/pexo-skills, 804 stars), Music (tadaspetra/loop, 296 stars), ElevenLabs Voiceover Generator (digitalsamba/claude-code-video-toolkit, 2.2k stars) and Sound Effects (tadaspetra/loop, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Elevenlabs Audio Prompting?

nodetool-ai (a GitHub organization) maintains it in nodetool-ai/nodetool, which has 560 GitHub stars. The repository holds 127 skills in this directory. The repository was last updated on October 11, 2026.

Source: nodetool-ai/nodetool on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.