Elevenlabs Core Workflow B
jeremylongshore/tons-of-skills-marketplace
Implement ElevenLabs speech-to-speech, sound effects, audio isolation, and speech-to-text.
Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction…
$ npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills audiopod --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/audiopod .claude/skills/audiopod && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "audiopod" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/audiopod into .claude/skills/audiopod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audiopod", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/audiopodType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills audiopod --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/audiopod .agents/skills/audiopod && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "audiopod" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/audiopod into .agents/skills/audiopod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audiopod", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills audiopod --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/audiopod .cursor/skills/audiopod && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "audiopod" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/audiopod into .cursor/skills/audiopod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audiopod", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LeoYeAI/openclaw-master-skills.git --path skills/audiopod--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills audiopod --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/audiopod .gemini/skills/audiopod && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "audiopod" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/audiopod into .gemini/skills/audiopod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audiopod", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LeoYeAI/openclaw-master-skills audiopodInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/audiopod .github/skills/audiopod && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "audiopod" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/audiopod into .github/skills/audiopod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audiopod", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LeoYeAI/openclaw-master-skills audiopod --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LeoYeAI/openclaw-master-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/audiopod .opencode/skills/audiopod && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "audiopod" agent skill from https://github.com/LeoYeAI/openclaw-master-skills/tree/main/skills/audiopod into .opencode/skills/audiopod/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "audiopod", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
audiopodUse AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction…
Audiopod is an agent skill from LeoYeAI/openclaw-master-skills. Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPODAPIKEY env var or pass…
Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `_meta.json`, `references/stems.md` and `references/tts.md`).
It sits in Media & Creative, covering Transcription, Music and audio generation and Text to speech and voice. It works with YouTube. The repository describes itself as: 🧠 Curated collection of 1209+ best OpenClaw skills — weekly updated by MyClaw.ai. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit e5199b5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlpipnpmffmpegFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.audiopod.aiyoutube.comAlso links to:
audiopod.aiFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
AUDIOPOD_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Audiopod loads about 5.5k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 137 tokens; SKILL.md has 648 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from LeoYeAI/openclaw-master-skills at commit e5199b5, republished under its MIT licence (© LeoYeAI). 648 words, ~5,505 tokens.
.claude/skills/audiopod/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Full audio processing API: music generation, stem separation, TTS, noise reduction, transcription, speaker separation, wallet management.
pip install audiopod # Python
npm install audiopod # Node.jsAuth: set AUDIOPOD_API_KEY env var or pass to client constructor.
ap_)from audiopod import AudioPod
client = AudioPod() # uses AUDIOPOD_API_KEY env var
# or: client = AudioPod(api_key="ap_...")Generate songs, rap, instrumentals, samples, and vocals from text prompts.
Tasks: text2music (song with vocals), text2rap (rap), prompt2instrumental (instrumental), lyric2vocals (vocals only), text2samples (loops/samples), audio2audio (style transfer), songbloom
# Generate a full song with lyrics
result = client.music.song(
prompt="Upbeat pop, synth, drums, 120 bpm, female vocals, radio-ready",
lyrics="Verse 1:\nWalking down the street on a sunny day\n\nChorus:\nWe're on fire tonight!",
duration=60
)
print(result["output_url"])
# Generate rap
result = client.music.rap(
prompt="Lo-Fi Hip Hop, 100 BPM, male rap, melancholy, keyboard chords",
lyrics="Verse 1:\nStarted from the bottom, now we climbing...",
duration=60
)
# Generate instrumental (no lyrics needed)
result = client.music.instrumental(
prompt="Atmospheric ambient soundscape, uplifting, driving mood",
duration=30
)
# Generic generate with explicit task
result = client.music.generate(
prompt="Electronic dance music, high energy",
task="text2samples", # any task type
duration=30
)
# Async: submit then poll
job = client.music.create(
prompt="Chill lofi beat",
duration=30,
task="prompt2instrumental"
)
result = client.music.wait_for_completion(job["id"], timeout=600)
# Get available genre presets
presets = client.music.get_presets()
# List/manage jobs
jobs = client.music.list(skip=0, limit=50)
job = client.music.get(job_id=123)
client.music.delete(job_id=123)# Song with lyrics
curl -X POST "https://api.audiopod.ai/api/v1/music/text2music" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"upbeat pop, synth, 120bpm, female vocals", "lyrics":"Walking down the street...", "audio_duration":60}'
# Rap
curl -X POST "https://api.audiopod.ai/api/v1/music/text2rap" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"Lo-Fi Hip Hop, male rap, 100 BPM", "lyrics":"Started from the bottom...", "audio_duration":60}'
# Instrumental
curl -X POST "https://api.audiopod.ai/api/v1/music/prompt2instrumental" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"ambient soundscape, uplifting", "audio_duration":30}'
# Samples/loops
curl -X POST "https://api.audiopod.ai/api/v1/music/text2samples" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"drum loop, sad mood", "audio_duration":15}'
# Vocals only
curl -X POST "https://api.audiopod.ai/api/v1/music/lyric2vocals" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"clean vocals, happy", "lyrics":"Eternal chorus of unity...", "audio_duration":30}'
# Check job status / get result
curl "https://api.audiopod.ai/api/v1/music/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Get genre presets
curl "https://api.audiopod.ai/api/v1/music/presets" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# List jobs
curl "https://api.audiopod.ai/api/v1/music/jobs?skip=0&limit=50" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Delete job
curl -X DELETE "https://api.audiopod.ai/api/v1/music/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"| Field | Required | Description |
|---|---|---|
| prompt | yes | Style/genre description |
| lyrics | for song/rap/vocals | Song lyrics with verse/chorus structure |
| audio_duration | no | Duration in seconds (default: 30) |
| genre_preset | no | Genre preset name (from presets endpoint) |
| display_name | no | Track display name |
Split audio into individual instrument/vocal tracks.
| Mode | Stems | Output | Use Case |
|---|---|---|---|
| single | 1 | Specified stem only | Vocal isolation, drum extraction |
| two | 2 | vocals + instrumental | Karaoke tracks |
| four | 4 | vocals, drums, bass, other | Standard remixing (default) |
| six | 6 | + guitar, piano | Full instrument separation |
| producer | 8 | + kick, snare, hihat | Beat production |
| studio | 12 | + cymbals, sub_bass, synth | Professional mixing |
| mastering | 16 | Maximum detail | Forensic analysis |
Single stem options: vocals, drums, bass, guitar, piano, other
# Sync: extract and wait for result
result = client.stems.separate(
url="https://youtube.com/watch?v=VIDEO_ID",
mode="six",
timeout=600
)
for stem, url in result["download_urls"].items():
print(f"{stem}: {url}")
# From local file
result = client.stems.separate(file="/path/to/song.mp3", mode="four")
# Single stem extraction
result = client.stems.separate(
url="https://youtube.com/watch?v=ID",
mode="single",
stem="vocals"
)
# Async: submit then poll
job = client.stems.extract(url="https://youtube.com/watch?v=ID", mode="six")
print(f"Job ID: {job['id']}")
status = client.stems.status(job["id"])
# or wait:
result = client.stems.wait_for_completion(job["id"], timeout=600)
# List available modes
modes = client.stems.modes()
# Job management
jobs = client.stems.list(skip=0, limit=50, status="COMPLETED")
job = client.stems.get(job_id=1234)
client.stems.delete(job_id=1234)# Extract from URL
curl -X POST "https://api.audiopod.ai/api/v1/stem-extraction/api/extract" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "url=https://youtube.com/watch?v=VIDEO_ID" \
-F "mode=six"
# Extract from file
curl -X POST "https://api.audiopod.ai/api/v1/stem-extraction/api/extract" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "file=@/path/to/song.mp3" \
-F "mode=four"
# Single stem
curl -X POST "https://api.audiopod.ai/api/v1/stem-extraction/api/extract" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "url=URL" \
-F "mode=single" \
-F "stem=vocals"
# Check job status
curl "https://api.audiopod.ai/api/v1/stem-extraction/status/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# List available modes
curl "https://api.audiopod.ai/api/v1/stem-extraction/modes" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# List jobs (filter by status: PENDING, PROCESSING, COMPLETED, FAILED)
curl "https://api.audiopod.ai/api/v1/stem-extraction/jobs?skip=0&limit=50&status=COMPLETED" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Get specific job
curl "https://api.audiopod.ai/api/v1/stem-extraction/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Delete job
curl -X DELETE "https://api.audiopod.ai/api/v1/stem-extraction/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"{
"id": 1234,
"status": "COMPLETED",
"download_urls": {
"vocals": "https://...",
"drums": "https://...",
"bass": "https://...",
"other": "https://..."
},
"quality_scores": {
"vocals": 0.95,
"drums": 0.88
}
}Generate speech from text with 50+ voices in 60+ languages. Supports voice cloning.
# Generate speech and wait for result
result = client.voice.generate(
text="Hello, world! This is a test.",
voice_id=123,
speed=1.0
)
print(result["output_url"])
# Async: submit then poll
job = client.voice.speak(
text="Hello world",
voice_id=123,
speed=1.0
)
status = client.voice.get_job(job["id"])
result = client.voice.wait_for_completion(job["id"], timeout=300)
# List all available voices
voices = client.voice.list()
for v in voices:
print(f"{v['id']}: {v['name']}")
# Clone a voice (needs ~5 sec audio sample)
new_voice = client.voice.create(
name="My Voice Clone",
audio_file="./sample.mp3",
description="Cloned from recording"
)
# Get/delete voice
voice = client.voice.get(voice_id=123)
client.voice.delete(voice_id=123)# List all voices
curl "https://api.audiopod.ai/api/v1/voice/voice-profiles" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Generate speech (FORM DATA, not JSON!)
curl -X POST "https://api.audiopod.ai/api/v1/voice/voices/{VOICE_UUID}/generate" \
-H "Authorization: Bearer $AUDIOPOD_API_KEY" \
-d "input_text=Hello world, this is a test" \
-d "audio_format=mp3" \
-d "speed=1.0"
# Poll job status
curl "https://api.audiopod.ai/api/v1/voice/tts-jobs/{JOB_ID}/status" \
-H "Authorization: Bearer $AUDIOPOD_API_KEY"
# SDK-style endpoints (alternative)
# Generate via SDK endpoint
curl -X POST "https://api.audiopod.ai/api/v1/voice/tts/generate" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Hello world","voice_id":123,"speed":1.0}'
# Poll via SDK endpoint
curl "https://api.audiopod.ai/api/v1/voice/tts/status/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# List voices (SDK endpoint)
curl "https://api.audiopod.ai/api/v1/voice/voices" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Clone a voice
curl -X POST "https://api.audiopod.ai/api/v1/voice/voices" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "name=My Voice" \
-F "file=@sample.mp3" \
-F "description=Cloned voice"
# Delete voice
curl -X DELETE "https://api.audiopod.ai/api/v1/voice/voices/VOICE_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"| Field | Required | Description |
|---|---|---|
| input_text | yes | Text to speak (max 5000 chars). Use input_text for raw HTTP, text for SDK |
| audio_format | no | mp3, wav, ogg (default: mp3) |
| speed | no | 0.25 - 4.0 (default: 1.0) |
| language | no | ISO code, auto-detected if omitted |
// Generate response
{"job_id": 12345, "status": "pending", "credits_reserved": 25}
// Status response (completed)
{"status": "completed", "output_url": "https://r2-url/generated.mp3"}input_text not text/api/v1/voice/tts/generate) uses JSON with field textffmpeg -i output.mp3 -c:a aac real.m4aSeparate audio by speaker with automatic diarization.
# Diarize and wait for result
result = client.speaker.identify(
file="./meeting.mp3",
num_speakers=3, # optional hint for accuracy
timeout=600
)
for segment in result["segments"]:
print(f"Speaker {segment['speaker']}: {segment['text']} [{segment['start']:.1f}s - {segment['end']:.1f}s]")
# From URL
result = client.speaker.identify(
url="https://youtube.com/watch?v=VIDEO_ID",
num_speakers=2
)
# Async: submit then poll
job = client.speaker.diarize(
file="./meeting.mp3",
num_speakers=3
)
result = client.speaker.wait_for_completion(job["id"], timeout=600)
# Job management
jobs = client.speaker.list(skip=0, limit=50, status="COMPLETED")
job = client.speaker.get(job_id=123)
client.speaker.delete(job_id=123)# Diarize from file
curl -X POST "https://api.audiopod.ai/api/v1/speaker/diarize" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "file=@meeting.mp3" \
-F "num_speakers=3"
# Diarize from URL
curl -X POST "https://api.audiopod.ai/api/v1/speaker/diarize" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "url=https://youtube.com/watch?v=VIDEO_ID" \
-F "num_speakers=2"
# Check job status
curl "https://api.audiopod.ai/api/v1/speaker/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# List jobs
curl "https://api.audiopod.ai/api/v1/speaker/jobs?skip=0&limit=50" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Delete job
curl -X DELETE "https://api.audiopod.ai/api/v1/speaker/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"Transcribe audio/video with speaker diarization, word-level timestamps, and multiple output formats.
# Transcribe URL and wait
result = client.transcription.transcribe(
url="https://youtube.com/watch?v=VIDEO_ID",
speaker_diarization=True,
min_speakers=2,
max_speakers=5,
timeout=600
)
print(f"Language: {result['detected_language']}")
for seg in result["segments"]:
print(f"[{seg['start']:.1f}s] {seg.get('speaker','?')}: {seg['text']}")
# Batch: multiple URLs at once
result = client.transcription.transcribe(
urls=["https://youtube.com/watch?v=ID1", "https://youtube.com/watch?v=ID2"],
speaker_diarization=True
)
# Upload local file
job = client.transcription.upload(
file_path="./recording.mp3",
language="en",
speaker_diarization=True
)
result = client.transcription.wait_for_completion(job["id"], timeout=600)
# Async: submit then poll
job = client.transcription.create(
url="https://youtube.com/watch?v=ID",
language="en",
speaker_diarization=True,
word_timestamps=True,
min_speakers=2,
max_speakers=4
)
result = client.transcription.wait_for_completion(job["id"], timeout=600)
# Get transcript in different formats
transcript_json = client.transcription.get_transcript(job_id=123, format="json")
transcript_srt = client.transcription.get_transcript(job_id=123, format="srt")
transcript_vtt = client.transcription.get_transcript(job_id=123, format="vtt")
transcript_txt = client.transcription.get_transcript(job_id=123, format="txt")
# Job management
jobs = client.transcription.list(skip=0, limit=50, status="COMPLETED")
job = client.transcription.get(job_id=123)
client.transcription.delete(job_id=123)# Transcribe from URL
curl -X POST "https://api.audiopod.ai/api/v1/transcribe/transcribe" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url":"https://youtube.com/watch?v=ID","enable_speaker_diarization":true,"word_timestamps":true}'
# Transcribe multiple URLs
curl -X POST "https://api.audiopod.ai/api/v1/transcribe/transcribe" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"urls":["URL1","URL2"],"enable_speaker_diarization":true}'
# Upload file for transcription
curl -X POST "https://api.audiopod.ai/api/v1/transcribe/transcribe-upload" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "files=@recording.mp3" \
-F "language=en" \
-F "enable_speaker_diarization=true"
# Get job status
curl "https://api.audiopod.ai/api/v1/transcribe/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Get transcript in specific format (json, srt, vtt, txt)
curl "https://api.audiopod.ai/api/v1/transcribe/jobs/JOB_ID/transcript?format=srt" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# List jobs
curl "https://api.audiopod.ai/api/v1/transcribe/jobs?offset=0&limit=50" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Delete job
curl -X DELETE "https://api.audiopod.ai/api/v1/transcribe/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"| Field | Required | Description |
|---|---|---|
| url / urls | yes (or file) | URL(s) to transcribe (YouTube, SoundCloud, direct links) |
| language | no | ISO 639-1 code (auto-detected if omitted) |
| enable_speaker_diarization | no | Enable speaker identification (default: false) |
| min_speakers / max_speakers | no | Speaker count hints for better diarization |
| word_timestamps | no | Enable word-level timestamps (default: true) |
Remove background noise from audio/video files.
# Denoise and wait for result
result = client.denoiser.denoise(file="./noisy-audio.mp3", timeout=600)
print(f"Clean audio: {result['output_url']}")
# From URL
result = client.denoiser.denoise(url="https://example.com/noisy.mp3")
# Async: submit then poll
job = client.denoiser.create(file="./noisy-audio.mp3")
result = client.denoiser.wait_for_completion(job["id"], timeout=600)
# From URL (async)
job = client.denoiser.create(url="https://example.com/noisy.mp3")
# Job management
jobs = client.denoiser.list(skip=0, limit=50, status="COMPLETED")
job = client.denoiser.get(job_id=123)
client.denoiser.delete(job_id=123)# Denoise from file
curl -X POST "https://api.audiopod.ai/api/v1/denoiser/denoise" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "file=@noisy-audio.mp3"
# Denoise from URL
curl -X POST "https://api.audiopod.ai/api/v1/denoiser/denoise" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-F "url=https://example.com/noisy.mp3"
# Check job status
curl "https://api.audiopod.ai/api/v1/denoiser/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# List jobs
curl "https://api.audiopod.ai/api/v1/denoiser/jobs?skip=0&limit=50" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Delete job
curl -X DELETE "https://api.audiopod.ai/api/v1/denoiser/jobs/JOB_ID" \
-H "X-API-Key: $AUDIOPOD_API_KEY"Check balance, estimate costs, and view usage history.
# Get current balance
balance = client.wallet.get_balance()
print(f"Balance: ${balance['balance_usd']}")
# Check if balance is sufficient for an operation
check = client.wallet.check_balance(
service_type="stem_extraction",
duration_seconds=180
)
print(f"Sufficient: {check['sufficient']}")
# Estimate cost before running
estimate = client.wallet.estimate_cost(
service_type="transcription",
duration_seconds=300
)
print(f"Cost: ${estimate['cost_usd']}")
# Get pricing for all services
pricing = client.wallet.get_pricing()
# View usage history
usage = client.wallet.get_usage(page=1, limit=50)# Get balance
curl "https://api.audiopod.ai/api/v1/api-wallet/balance" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Check balance sufficiency
curl -X POST "https://api.audiopod.ai/api/v1/api-wallet/check-balance" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"service_type":"stem_extraction","duration_seconds":180}'
# Estimate cost
curl -X POST "https://api.audiopod.ai/api/v1/api-wallet/estimate-cost" \
-H "X-API-Key: $AUDIOPOD_API_KEY" \
-H "Content-Type: application/json" \
-d '{"service_type":"transcription","duration_seconds":300}'
# Get pricing
curl "https://api.audiopod.ai/api/v1/api-wallet/pricing" \
-H "X-API-Key: $AUDIOPOD_API_KEY"
# Usage history
curl "https://api.audiopod.ai/api/v1/api-wallet/usage?page=1&limit=50" \
-H "X-API-Key: $AUDIOPOD_API_KEY"| Service | Endpoint | Method |
|---|---|---|
| Music | /api/v1/music/{task} | POST |
| Music jobs | /api/v1/music/jobs/{id} | GET/DELETE |
| Music presets | /api/v1/music/presets | GET |
| Stems | /api/v1/stem-extraction/api/extract | POST (multipart) |
| Stems status | /api/v1/stem-extraction/status/{id} | GET |
| Stems modes | /api/v1/stem-extraction/modes | GET |
| Stems jobs | /api/v1/stem-extraction/jobs | GET |
| TTS generate | /api/v1/voice/voices/{uuid}/generate | POST (form data) |
| TTS generate (SDK) | /api/v1/voice/tts/generate | POST (JSON) |
| TTS status | /api/v1/voice/tts-jobs/{id}/status | GET |
| TTS status (SDK) | /api/v1/voice/tts/status/{id} | GET |
| Voice list | /api/v1/voice/voice-profiles | GET |
| Voice list (SDK) | /api/v1/voice/voices | GET |
| Speaker | /api/v1/speaker/diarize | POST (multipart) |
| Speaker jobs | /api/v1/speaker/jobs/{id} | GET/DELETE |
| Transcribe URL | /api/v1/transcribe/transcribe | POST (JSON) |
| Transcribe upload | /api/v1/transcribe/transcribe-upload | POST (multipart) |
| Transcript output | /api/v1/transcribe/jobs/{id}/transcript?format= | GET |
| Transcribe jobs | /api/v1/transcribe/jobs | GET |
| Denoise | /api/v1/denoiser/denoise | POST (multipart) |
| Denoise jobs | /api/v1/denoiser/jobs/{id} | GET/DELETE |
| Wallet balance | /api/v1/api-wallet/balance | GET |
| Wallet pricing | /api/v1/api-wallet/pricing | GET |
| Wallet usage | /api/v1/api-wallet/usage | GET |
Two auth styles work:
X-API-Key: ap_... — works for most endpointsAuthorization: Bearer ap_... — works for TTS generate/statusoutput_url in job status© LeoYeAI, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/audiopod of LeoYeAI/openclaw-master-skills.
Open the folder on GitHubat commit e5199b5
Audiopod next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Audiopod this skillLeoYeAI/openclaw-master-skills | 2.2k | — | ~5.5k | Automated safety check: Pass | MIT | |
| Elevenlabs Core Workflow Bjeremylongshore/tons-of-skills-marketplace | 2.8k | — | ~1.6k | Automated safety check: Pass | MIT | |
| HyperFrames Media Useheygen-com/hyperframes | 60k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| Elevenlabszapier/connectors | 177 | — | ~3.7k | Automated safety check: Pass | Elastic-2.0 | |
| Speech Buildcnemri/google-genai-skills | 127 | — | ~430 | Automated safety check: Pass | MIT | |
| Voice AI Engine Developmentaiskillstore/marketplace | 433 | 5 repos | ~5.7k | Automated safety check: Pass | None |
jeremylongshore/tons-of-skills-marketplace
Implement ElevenLabs speech-to-speech, sound effects, audio isolation, and speech-to-text.
heygen-com/hyperframes
Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.
zapier/connectors
Agent-callable ElevenLabs tools — generate spoken audio from text, create sound effects and multi-speaker dialogue, re-voice and clean up audio, transcribe audio and video, design synthetic voices…
cnemri/google-genai-skills
Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models.
aiskillstore/marketplace
Build real-time conversational AI voice engines using async worker pipelines, streaming transcription, LLM agents, and TTS synthesis with interrupt handling and multi-provider support
daymade/claude-code-skills
Transcribes Chinese/English audio with StepFun's stepaudio-3-asr-max via its SSE endpoint (not /v1/audio/transcriptions) — one call handles long-form audio with no chunking.
LeoYeAI/openclaw-master-skills
Manages pipelines on a DevOps quality and efficiency platform through its OpenAPI: list workspaces and templates, create, update, run and cancel pipelines, and read run records.
LeoYeAI/openclaw-master-skills
Patches OpenClaw's Feishu extension so an edited document triggers an isolated agent session that reads the doc and replies inline, turning it into a live chat space.
LeoYeAI/openclaw-master-skills
Multi-context memory management system for OpenClaw agents with group-isolated storage, global shared memory, workspace organization, and group-specific skills isolation.
LeoYeAI/openclaw-master-skills
Runs a brand's AI-search visibility work end to end: diagnosing how AI platforms represent it, repositioning it, producing AI-optimized content and monitoring ongoing mentions.
LeoYeAI/openclaw-master-skills
Installs and authenticates the gws CLI, then automates Gmail, Drive, Sheets, Calendar, Docs, Chat and Tasks with ready-made recipes, persona bundles and security audits.
LeoYeAI/openclaw-master-skills
Runs four advisor roles, a fitness coach, nutritionist, data analyst and TCM practitioner, to build a health profile and track workouts, diet and wellness over time.
Works with
Categories
Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction…. Audiopod is an agent skill from LeoYeAI/openclaw-master-skills. Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction.
Audiopod fits situations like: the user needs to generate music/songs/rap from text; split a song into stems/vocals/instruments; generate speech from text; clean up noisy audio.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a claude-code`. Or copy the skill folder (skills/audiopod in LeoYeAI/openclaw-master-skills) into .claude/skills/audiopod in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a codex`. Or copy the skill folder (skills/audiopod in LeoYeAI/openclaw-master-skills) into .agents/skills/audiopod in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LeoYeAI/openclaw-master-skills --skill audiopod -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audiopod, .gemini/skills/audiopod, .github/skills/audiopod and .opencode/skills/audiopod in your project.
Going by SKILL.md and its folder, Audiopod needs the command-line tools its instructions call (curl, pip, npm and ffmpeg) and credentials named AUDIOPOD_API_KEY. Our summary lists: Python 3; Node.js; A credential in AUDIOPOD_API_KEY.
SKILL.md names 3 domains. In commands or code: api.audiopod.ai and youtube.com; the agent is likely to contact these when it follows the instructions. As links in the text: audiopod.ai. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Audiopod is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 728 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Audiopod: Elevenlabs Core Workflow B (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Elevenlabs (zapier/connectors, 177 stars) and Speech Build (cnemri/google-genai-skills, 127 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LeoYeAI (a GitHub user) maintains it in LeoYeAI/openclaw-master-skills, which has 2,161 GitHub stars. The repository holds 1,235 skills in this directory. The repository was last updated on July 20, 2026.
Source: LeoYeAI/openclaw-master-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.