Agent skill

Videoagent Audio Studio

by pexoai in pexoai/pexo-skills

Tired of juggling multiple audio APIs?. An agent skill from pexoai/pexo-skills.

MITAuto-check passedMedia & Creative

Install Videoagent Audio Studio

skills CLI
$ npx skills add pexoai/pexo-skills --skill videoagent-audio-studio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pexoai/pexo-skills videoagent-audio-studio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pexoai/pexo-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/videoagent-audio-studio .claude/skills/videoagent-audio-studio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
videoagent-audio-studio
GitHub stars
804
Token cost
~1.7k tokens
SKILL.md length
532 words
Files
10
Skills in repo
18
Repo updated
First seen
Licence
MIT

At a glance

Tired of juggling multiple audio APIs?. An agent skill from pexoai/pexo-skills.

  • Works in 2 steps: Start the AudioMind server (once per… → Route the request
  • You want to generate any audio without managing multiple API keys
  • SKILL.md covers Quick Reference, How to Use, Example Conversations and Setup, plus 3 more sections
  • Runs JavaScript and Shell scripts from its folder; calls bash, npm and vercel; needs ELEVENLABS_API_KEY and FAL_KEY

What it does

Videoagent Audio Studio is an agent skill from pexoai/pexo-skills. Tired of juggling multiple audio APIs? This skill gives you one-command access to TTS, music generation, sound effects, and voice cloning. Use when you want to generate any audio without managing multiple API keys.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files (for example `cli.js`, `proxy/api/audio.js` and `proxy/api/stats.js`).

It sits in Media & Creative, covering Text to speech and voice and Music and audio generation. It works with fal, Vercel and ElevenLabs. The repository describes itself as: A collection of open-source Agent Skills for content creation — images, audio, and video. The licence is MIT.

When your agent uses it

  • You want to generate any audio without managing multiple API keys
  • Tasks that involve Text to speech and voice
  • Tasks that involve Music and audio generation

Example prompts

  • “/videoagent-audio-studio”

Requirements

  • Node.js
  • A Bash shell
  • A credential in ELEVENLABS_API_KEY
  • A credential in FAL_KEY

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Start the AudioMind server (once per session)
  2. Route the request

What it can do on your machine

Read from SKILL.md and the folder at commit f724267. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (JavaScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • npm
    • vercel

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • elevenlabs.io
    • fal.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELEVENLABS_API_KEY
    • FAL_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Videoagent Audio Studio loads about 1.7k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 532 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pexoai/pexo-skills at commit f724267, republished under its MIT licence (© pexoai). 532 words, ~1,715 tokens.

Download SKILL.mdSave it as .claude/skills/videoagent-audio-studio/SKILL.md (or your agent's skills folder). This skill also uses 9 other files; get the full folder from GitHub.
name
videoagent-audio-studio
description
Tired of juggling multiple audio APIs? This skill gives you one-command access to TTS, music generation, sound effects, and voice cloning. Use when you want to generate any audio without managing multiple API keys.
version
3.0.0
author
wells
emoji
🎙️
tags
video, audio, tts, music, sfx, voice-clone, elevenlabs, fal
homepage
https://github.com/pexoai/audiomind-skill

🎙️ VideoAgent Audio Studio

Use when: User asks to generate speech, narrate text, create a voice-over, compose music, or produce a sound effect.

VideoAgent Audio Studio is a smart audio dispatcher. It analyzes your request and routes it to the best available model — ElevenLabs for speech and music, fal.ai for fast SFX — and returns a ready-to-use audio URL.


Quick Reference

Request TypeBest ModelLatency
Narrate text / Voice-overelevenlabs-tts-v3~3s
Low-latency TTS (real-time)elevenlabs-tts-turbo<1s
Background musiccassetteai-music~15s
Sound effectelevenlabs-sfx~5s
Clone a voice from audioelevenlabs-voice-clone~10s

How to Use

1. Start the AudioMind server (once per session)
bash
bash {baseDir}/tools/start_server.sh

This starts the ElevenLabs MCP server on port 8124. The skill uses it for all audio generation.

2. Route the request

Analyze the user's request and call the appropriate tool via the MCP server:

Text-to-Speech (TTS)

When user asks to "narrate", "read aloud", "say", or "create a voice-over":

Use MCP tool: text_to_speech
  text: "<the text to narrate>"
  voice_id: "JBFqnCBsd6RMkjVDRZzb"   # Default: "George" (professional, neutral)
  model_id: "eleven_multilingual_v2"   # Use "eleven_turbo_v2_5" for low latency

Music Generation

When user asks to "compose", "create background music", or "make a soundtrack":

Use MCP tool: text_to_sound_effects  (via cassetteai-music on fal.ai)
  prompt: "<music description, e.g. 'upbeat lo-fi hip hop, 90 seconds'>"
  duration_seconds: <duration>

Sound Effect (SFX)

When user asks for a specific sound (e.g., "a door creaking", "rain on a window"):

Use MCP tool: text_to_sound_effects
  text: "<sound description>"
  duration_seconds: <1-22>

Voice Cloning

When user provides an audio sample and wants to clone the voice:

Use MCP tool: voice_add
  name: "<voice name>"
  files: ["<audio_file_url>"]

Example Conversations

User: "Voice this text for me: Welcome to our product launch"

→ Route to: text_to_speech
  text: "Welcome to our product launch"
  voice_id: "JBFqnCBsd6RMkjVDRZzb"
  model_id: "eleven_multilingual_v2"

🎙️ Voiceover done! Listen here


User: "Generate 60 seconds of relaxing background music for a podcast"

→ Route to: cassetteai-music (fal.ai)
  prompt: "relaxing lo-fi background music for a podcast, gentle piano and soft beats, 60 seconds"
  duration_seconds: 60

🎵 Background music ready! Listen here


User: "Generate a sci-fi style door opening sound effect"

→ Route to: text_to_sound_effects
  text: "a futuristic sci-fi door sliding open with a hydraulic hiss"
  duration_seconds: 3

Setup

Required

Set ELEVENLABS_API_KEY in ~/.openclaw/openclaw.json:

json
{
  "skills": {
    "entries": {
      "videoagent-audio-studio": {
        "enabled": true,
        "env": {
          "ELEVENLABS_API_KEY": "your_elevenlabs_key_here"
        }
      }
    }
  }
}

Get your key at elevenlabs.io/app/settings/api-keys.

Optional (for fal.ai music & SFX models)
json
"FAL_KEY": "your_fal_key_here"

Get your key at fal.ai/dashboard/keys.


Self-Hosting the Proxy

The cli.js connects to a hosted proxy by default. If you want full control — or need to serve users in regions where vercel.app is blocked — you can deploy your own instance from the proxy/ directory.

Quick Deploy (Vercel)
bash
cd proxy
npm install
vercel --prod
Environment Variables

Set these in your Vercel project (Dashboard → Settings → Environment Variables):

VariableRequired ForWhere to Get
ELEVENLABS_API_KEYTTS, SFX, Voice Cloneelevenlabs.io/app/settings/api-keys
FAL_KEYMusic generationfal.ai/dashboard/keys
VALID_PRO_KEYS(Optional) Restrict accessComma-separated list of allowed client keys
Show full SKILL.md (195 more words)Show less
Point cli.js to Your Proxy
bash
export AUDIOMIND_PROXY_URL="https://your-domain.com/api/audio"

Or set it in ~/.openclaw/openclaw.json:

json
{
  "skills": {
    "entries": {
      "videoagent-audio-studio": {
        "env": {
          "AUDIOMIND_PROXY_URL": "https://your-domain.com/api/audio"
        }
      }
    }
  }
}

If your users are in mainland China, bind a custom domain in Vercel Dashboard → Settings → Domains to avoid DNS issues with vercel.app.


Model Reference

Model IDTypeProviderNotes
eleven_multilingual_v2TTSElevenLabsBest quality, supports 29 languages
eleven_turbo_v2_5TTSElevenLabsUltra-low latency, ideal for real-time
eleven_monolingual_v1TTSElevenLabsEnglish only, fastest
cassetteai-musicMusicfal.aiReliable, fast music generation
elevenlabs-sfxSFXElevenLabsHigh-quality sound effects (up to 22s)
elevenlabs-voice-cloneCloneElevenLabsClone any voice from a short audio sample

Changelog

v3.0.0
  • Simplified routing table: Removed unstable/offline models from the main reference. The skill now only surfaces models that reliably work.
  • Clearer use-case triggers: Added "Use when" section so the agent activates this skill at the right moment.
  • Unified setup: Single ELEVENLABS_API_KEY is all you need to get started. FAL_KEY is now optional.
  • Removed polling complexity: Music generation now uses cassetteai-music by default, which completes synchronously.
v2.1.0
  • Added async workflow for long-running music generation tasks.
  • Added cassetteai-music as a stable alternative for music generation.
v2.0.0
  • Migrated to ElevenLabs MCP server architecture.
  • Added voice cloning support.
v1.0.0
  • Initial release with TTS, music, and SFX routing.

© pexoai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 9 other files in skills/videoagent-audio-studio of pexoai/pexo-skills.

  • SKILL.md
  • cli.js
  • proxy/.gitignore
  • proxy/api/audio.js
  • proxy/api/stats.js
  • proxy/package-lock.json
  • proxy/package.json
  • proxy/usage-store.js
  • proxy/vercel.json
  • tools/start_server.sh

Open the folder on GitHubat commit f724267

Compare with similar skills

Videoagent Audio Studio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Videoagent Audio Studio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Videoagent Audio Studio this skillpexoai/pexo-skills804—~1.7kAutomated safety check: PassMIT
Elevenlabs Audio Promptingnodetool-ai/nodetool560—~1.9kAutomated safety check: PassAGPL-3.0
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT
ElevenLabs Voiceover Generatordigitalsamba/claude-code-video-toolkit2.2k1 repos~2.7kAutomated safety check: NotesMIT
Sound Effectstadaspetra/loop2962 repos~1.1kAutomated safety check: PassMIT
Hyperframes Mediachmonitor/chmonitor3001 repos~2.8kAutomated safety check: NotesGPL-3.0

Similar skills

  • Elevenlabs Audio Prompting

    nodetool-ai/nodetool

    Direct ElevenLabs speech, dialogue, sound effects and music — the bracketed audio tags v3 acts on and why the voice decides whether a tag lands, stability as the delivery dial, punctuation instead…

    560 GitHub stars~1.9k tokensUpdated today
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • ElevenLabs Voiceover Generator

    digitalsamba/claude-code-video-toolkit

    Generates narration, sound effects and cloned voices through the ElevenLabs API, with model and setting choices tuned to the content's style.

    2.2k GitHub starsUsed in 1 repo~2.7k tokens
    Media & CreativeAuto-check: notes
  • Sound Effects

    tadaspetra/loop

    Generate sound effects from text descriptions using ElevenLabs.

    296 GitHub starsUsed in 2 repos~1.1k tokens
    Media & CreativeAuto-check passed
  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    300 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • Fal AI

    mikeOnBreeze/cc-crossbeam

    This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API.

    293 GitHub stars~1.9k tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes

More from pexoai/pexo-skills

All 18 skills in this repo
  • AI Video Generation

    pexoai/pexo-skills

    Generate AI video from any input — text, image, or script — with Pexo.

    804 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Explainer Video

    pexoai/pexo-skills

    Create an explainer video with narration using Pexo. An agent skill from pexoai/pexo-skills.

    804 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Founder Video

    pexoai/pexo-skills

    Make a founder video with Pexo — built for solo founders and small teams.

    804 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Image To Video

    pexoai/pexo-skills

    Animate a still image into a finished, moving video with Pexo.

    804 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Launch Video

    pexoai/pexo-skills

    Make a launch video for your startup or product with Pexo. An agent skill from pexoai/pexo-skills.

    804 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Make A Video

    pexoai/pexo-skills

    Make a complete video from a simple idea with Pexo. An agent skill from pexoai/pexo-skills.

    804 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Videoagent Audio Studio

What does Videoagent Audio Studio do?

Tired of juggling multiple audio APIs?. An agent skill from pexoai/pexo-skills. Videoagent Audio Studio is an agent skill from pexoai/pexo-skills. Tired of juggling multiple audio APIs?

When should I use Videoagent Audio Studio?

Videoagent Audio Studio fits situations like: you want to generate any audio without managing multiple API keys; tasks that involve Text to speech and voice; tasks that involve Music and audio generation.

How do I install Videoagent Audio Studio in Claude Code?

Run `npx skills add pexoai/pexo-skills --skill videoagent-audio-studio -a claude-code`. Or copy the skill folder (skills/videoagent-audio-studio in pexoai/pexo-skills) into .claude/skills/videoagent-audio-studio in your project. Claude Code loads it when a task matches its description.

How do I install Videoagent Audio Studio in Codex?

Run `npx skills add pexoai/pexo-skills --skill videoagent-audio-studio -a codex`. Or copy the skill folder (skills/videoagent-audio-studio in pexoai/pexo-skills) into .agents/skills/videoagent-audio-studio in your project. Codex loads it when a task matches its description.

Can I use Videoagent Audio Studio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pexoai/pexo-skills --skill videoagent-audio-studio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/videoagent-audio-studio, .gemini/skills/videoagent-audio-studio, .github/skills/videoagent-audio-studio and .opencode/skills/videoagent-audio-studio in your project.

What does Videoagent Audio Studio need to run?

Going by SKILL.md and its folder, Videoagent Audio Studio needs JavaScript and a shell for the scripts in its folder, the command-line tools its instructions call (bash, npm and vercel) and credentials named ELEVENLABS_API_KEY and FAL_KEY. Our summary lists: Node.js; A Bash shell; A credential in ELEVENLABS_API_KEY; A credential in FAL_KEY.

Does Videoagent Audio Studio access the network?

SKILL.md names 2 domains. As links in the text: elevenlabs.io and fal.ai. This is read from the text; nothing was executed.

Is Videoagent Audio Studio safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Videoagent Audio Studio use?

Videoagent Audio Studio is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Videoagent Audio Studio use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Videoagent Audio Studio?

Skills that share tags, products or a category with Videoagent Audio Studio: Elevenlabs Audio Prompting (nodetool-ai/nodetool, 560 stars), Music (tadaspetra/loop, 296 stars), ElevenLabs Voiceover Generator (digitalsamba/claude-code-video-toolkit, 2.2k stars) and Sound Effects (tadaspetra/loop, 296 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Videoagent Audio Studio?

pexoai (a GitHub user) maintains it in pexoai/pexo-skills, which has 804 GitHub stars. The repository holds 18 skills in this directory. The repository was last updated on August 20, 2026.

Source: pexoai/pexo-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.