Audio Extraction: Extracting audio from videos, converting formats, and managing audio collections

MITAuto-check passedMedia & Creative

Install Audio Extraction

skills CLI
$ npx skills add cosmicstack-labs/mercury-agent-skills --skill audio-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install cosmicstack-labs/mercury-agent-skills audio-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/cosmicstack-labs/mercury-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/categories/media-download/audio-extraction .claude/skills/audio-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio-extraction
GitHub stars
476
Token cost
~4.7k tokens
SKILL.md length
618 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Audio Extraction: Extracting audio from videos, converting formats, and managing audio collections

  • Works in 4 steps: Source Quality Determines Output Quality → Choose the Right Format for the Use Case → Metadata Is Not Optional → …
  • Media & Creative work in your project
  • SKILL.md covers Core Principles, Audio Extraction with yt-dlp, FFmpeg Audio Processing and Metadata Tagging, plus 4 more sections
  • Calls yt-dlp, ffmpeg and pip; reaches youtube.com

What it does

Audio Extraction is an agent skill from cosmicstack-labs/mercury-agent-skills. Audio Extraction: Extracting audio from videos, converting formats, and managing audio collections

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative. The repository describes itself as: A curated registry of reusable Mercury Agent, Open Claw or Hermes Agent skills designed for real developer workflows, persistent memory, and token-efficient execution. The licence is MIT.

When your agent uses it

  • Media & Creative work in your project

Example prompts

  • “/audio-extraction”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Source Quality Determines Output Quality
  2. Choose the Right Format for the Use Case
  3. Metadata Is Not Optional
  4. Preserve the Original

What it can do on your machine

Read from SKILL.md and the folder at commit 30392fb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • yt-dlp
    • ffmpeg
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio Extraction loads about 4.7k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 618 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from cosmicstack-labs/mercury-agent-skills at commit 30392fb, republished under its MIT licence (© cosmicstack-labs). 618 words, ~4,688 tokens.

Download SKILL.mdSave it as .claude/skills/audio-extraction/SKILL.md (or your agent's skills folder).
name
audio-extraction
description
Audio Extraction: Extracting audio from videos, converting formats, and managing audio collections
metadata.author
cosmicstack-labs
metadata.version
1.0.0
metadata.category
media-download
metadata.tags
audio-extraction, mp3-conversion, audio-conversion, ffmpeg, music

Audio Extraction

Extract high-quality audio from video files, convert between formats, manage metadata, and build organized audio collections. This skill covers everything from one-off audio rips to batch processing pipelines.

Core Principles

1. Source Quality Determines Output Quality

You cannot create quality that wasn't captured. Start with the highest quality source available — lossy-to-lossy transcoding degrades audio further. Always extract from the best original source.

2. Choose the Right Format for the Use Case
  • MP3 (lossy): Universal compatibility, great for music players and portable devices
  • FLAC (lossless): Archival quality, for listening on quality equipment or future transcoding
  • AAC/M4A: Better quality than MP3 at the same bitrate, native to Apple ecosystem
  • OGG/Opus: Best quality-per-bitrate, perfect for streaming and podcasts
  • WAV (uncompressed): Editing and production, not for everyday listening
3. Metadata Is Not Optional

Untagged audio files are unmanageable at scale. Proper ID3 tags, cover art, and consistent naming conventions turn a pile of files into a browsable music library.

4. Preserve the Original

Always keep a copy of the original file or at minimum log what source was used. Once you transcode, you lose information. Archival means keeping the best available original plus a convenient playback copy.


Audio Extraction with yt-dlp

Basic Audio Extraction
bash
# Simplest audio extraction (best quality)
yt-dlp -x "https://youtube.com/watch?v=VIDEO_ID"

# Specific audio format
yt-dlp -x --audio-format mp3 "https://youtube.com/watch?v=VIDEO_ID"

# Best quality with metadata
yt-dlp -x --audio-format mp3 --audio-quality 0 \
  --embed-thumbnail --embed-metadata "URL"
Format Conversion Options
bash
# MP3 at various quality levels
yt-dlp -x --audio-format mp3 --audio-quality 0 "URL"     # 320kbps (best)
yt-dlp -x --audio-format mp3 --audio-quality 2 "URL"     # ~256kbps
yt-dlp -x --audio-format mp3 --audio-quality 5 "URL"     # ~192kbps (good)
yt-dlp -x --audio-format mp3 --audio-quality 9 "URL"     # ~128kbps (acceptable)

# FLAC (lossless)
yt-dlp -x --audio-format flac --audio-quality 0 "URL"

# AAC/M4A
yt-dlp -x --audio-format m4a "URL"

# Opus (best quality-per-bitrate)
yt-dlp -x --audio-format opus "URL"

# WAV (uncompressed)
yt-dlp -x --audio-format wav "URL"
Audio-Only Format Selection
bash
# List available audio formats
yt-dlp -F "URL" | grep -E "audio|opus|aac|mp3|m4a"

# Download specific audio stream
yt-dlp -f "140" "URL"  # 128kbps AAC (YouTube standard)

# Download highest bitrate audio
yt-dlp -f "bestaudio[abr>128]/bestaudio" "URL"

# Download Opus stream (YouTube music)
yt-dlp -f "251" "URL"  # 160kbps Opus

FFmpeg Audio Processing

Format Conversion
bash
# Convert MP4 to MP3
ffmpeg -i input.mp4 -vn -acodec libmp3lame -ab 320k output.mp3

# Convert any video to FLAC
ffmpeg -i input.mkv -vn -c:a flac output.flac

# Batch convert all MP4s in directory
for f in *.mp4; do
  ffmpeg -i "$f" -vn -acodec libmp3lame -ab 320k "${f%.mp4}.mp3"
done
Trimming Audio
bash
# Trim from 30s to 1m30s
ffmpeg -i input.mp3 -ss 00:00:30 -to 00:01:30 -c copy output.mp3

# Trim from start for 45 seconds
ffmpeg -i input.mp3 -t 45 -c copy output.mp3

# Trim with re-encoding (for precise cuts)
ffmpeg -i input.mp3 -ss 00:00:30 -to 00:01:30 output.mp3
Merging Audio Files
bash
# Concatenate with ffmpeg (same format)
ffmpeg -i "concat:file1.mp3|file2.mp3|file3.mp3" -c copy merged.mp3

# Using concat demuxer
echo "file 'part1.mp3'" > files.txt
echo "file 'part2.mp3'" >> files.txt
echo "file 'part3.mp3'" >> files.txt
ffmpeg -f concat -safe 0 -i files.txt -c copy merged.mp3

# Merge with crossfade
ffmpeg -i part1.mp3 -i part2.mp3 -filter_complex \
  "[0:a][1:a]acrossfade=d=2:c1=tri:c2=tri[a]" \
  -map "[a]" merged.mp3
Audio Normalization
bash
# EBU R128 loudness normalization (broadcast standard)
ffmpeg -i input.mp3 -af loudnorm=I=-16:LRA=11:TP=-1.5 output.mp3

# Peak normalization (simpler)
ffmpeg -i input.mp3 -af volume=3dB output.mp3

# Dynamic range compression
ffmpeg -i input.mp3 -af acompressor=threshold=-21dB:ratio=9:attack=200:release=1000 output.mp3

# Normalize batch files
for f in *.mp3; do
  ffmpeg -i "$f" -af loudnorm=I=-16:LRA=11:TP=-1.5 "normalized_$f"
done

Metadata Tagging

Using eyeD3 (MP3)
bash
# Install eyeD3
pip install eyeD3

# Set basic tags
eyeD3 -a "Artist Name" -A "Album Title" -t "Song Title" -n 1 -N 10 track.mp3

# Set genre and year
eyeD3 -G "Rock" -Y 2024 track.mp3

# Add album art
eyeD3 --add-image cover.jpg:FRONT_COVER track.mp3

# Remove all tags
eyeD3 --remove-all track.mp3
Using mutagen (Python - All Formats)
python
from mutagen.mp3 import MP3
from mutagen.id3 import ID3, TIT2, TPE1, TALB, TRCK, TYER, APIC
import os

def tag_audio_file(filepath, metadata, cover_art_path=None):
    """
    Tag an audio file with comprehensive metadata.
    
    Args:
        filepath: Path to the audio file
        metadata: Dict with keys: title, artist, album, track, year, genre
        cover_art_path: Path to cover art image
    """
    audio = MP3(filepath, ID3=ID3)
    audio.tags.add(TIT2(encoding=3, text=metadata['title']))
    audio.tags.add(TPE1(encoding=3, text=metadata['artist']))
    audio.tags.add(TALB(encoding=3, text=metadata['album']))
    audio.tags.add(TRCK(encoding=3, text=str(metadata['track'])))
    audio.tags.add(TYER(encoding=3, text=str(metadata['year'])))
    
    if cover_art_path and os.path.exists(cover_art_path):
        with open(cover_art_path, 'rb') as img:
            audio.tags.add(
                APIC(
                    encoding=3,
                    mime='image/jpeg',
                    type=3,  # Front cover
                    desc='Cover',
                    data=img.read()
                )
            )
    
    audio.save()

# Usage
tag_audio_file('track.mp3', {
    'title': 'Bohemian Rhapsody',
    'artist': 'Queen',
    'album': 'A Night at the Opera',
    'track': 11,
    'year': 1975,
    'genre': 'Rock'
}, 'cover.jpg')
Batch Metadata from Filename
python
import os
import re
from mutagen.mp3 import MP3
from mutagen.id3 import ID3, TIT2, TPE1, TALB

def tag_from_filename(directory, pattern=r"(.+?) - (.+?) - (.+)\.mp3"):
    """
    Tag files based on filename pattern.
    Default pattern: "Artist - Album - Title.mp3"
    """
    for filename in os.listdir(directory):
        if not filename.endswith('.mp3'):
            continue
        
        match = re.match(pattern, filename)
        if not match:
            continue
        
        artist, album, title = match.groups()
        filepath = os.path.join(directory, filename)
        
        audio = MP3(filepath, ID3=ID3)
        audio.tags.add(TPE1(encoding=3, text=artist.strip()))
        audio.tags.add(TALB(encoding=3, text=album.strip()))
        audio.tags.add(TIT2(encoding=3, text=title.strip()))
        audio.save()
        
        print(f"Tagged: {filename} → {artist} / {album} / {title}")

# Usage
tag_from_filename("~/Music/Downloads/")

Podcast RSS Feed Downloads

Using yt-dlp for Podcasts
bash
# Download podcast episode from RSS
yt-dlp -x --audio-format mp3 --audio-quality 0 "PODCAST_RSS_URL"

# Download only the latest episode
yt-dlp --playlist-end 1 -x --audio-format mp3 "RSS_URL"

# Download with consistent naming
yt-dlp -o "%(title)s.%(ext)s" -x --audio-format mp3 "RSS_URL"
Using gPodder (CLI)
bash
# Install gPodder
pip install gpodder

# Subscribe to a podcast
gpo add "https://example.com/podcast/rss"

# Download new episodes
gpo download

# List subscriptions
gpo list
Custom Podcast Downloader
python
import feedparser
import requests
import os
from urllib.parse import urlparse

def download_podcast_episodes(rss_url, output_dir="~/Podcasts"):
    """Download all episodes from an RSS feed."""
    output_dir = os.path.expanduser(output_dir)
    os.makedirs(output_dir, exist_ok=True)
    
    feed = feedparser.parse(rss_url)
    podcast_title = feed.feed.get('title', 'Unknown Podcast')
    podcast_dir = os.path.join(output_dir, podcast_title)
    os.makedirs(podcast_dir, exist_ok=True)
    
    for entry in feed.entries:
        title = entry.get('title', 'Unknown Episode')
        
        # Sanitize filename
        safe_title = "".join(c for c in title if c.isalnum() or c in ' -_').rstrip()
        
        # Find audio enclosure
        for link in entry.get('links', []):
            if link.get('type', '').startswith('audio/'):
                audio_url = link['href']
                ext = os.path.splitext(urlparse(audio_url).path)[1] or '.mp3'
                filepath = os.path.join(podcast_dir, f"{safe_title}{ext}")
                
                if os.path.exists(filepath):
                    print(f"✓ Already downloaded: {title}")
                    continue
                
                print(f"↓ Downloading: {title}")
                response = requests.get(audio_url, stream=True)
                with open(filepath, 'wb') as f:
                    for chunk in response.iter_content(chunk_size=8192):
                        if chunk:
                            f.write(chunk)
                print(f"✓ Saved: {filepath}")
                break

# Usage
download_podcast_episodes("https://feeds.example.com/podcast/rss.xml")

Batch Audio Extraction

Process Multiple Files
bash
# Extract audio from all videos in directory
for f in *.mp4 *.mkv *.webm; do
  [ -e "$f" ] || continue
  ffmpeg -i "$f" -vn -acodec libmp3lame -ab 320k "${f%.*}.mp3"
done
Recursive Directory Processing
python
import os
import subprocess

def extract_audio_recursive(root_dir, output_format='mp3', bitrate='320k'):
    """Extract audio from all video files in directory tree."""
    video_extensions = {'.mp4', '.mkv', '.webm', '.avi', '.mov', '.flv'}
    
    for dirpath, dirnames, filenames in os.walk(root_dir):
        for filename in filenames:
            ext = os.path.splitext(filename)[1].lower()
            if ext not in video_extensions:
                continue
            
            input_path = os.path.join(dirpath, filename)
            output_name = os.path.splitext(filename)[0] + f'.{output_format}'
            output_path = os.path.join(dirpath, output_name)
            
            if os.path.exists(output_path):
                print(f"✓ Already exists: {output_name}")
                continue
            
            print(f"⟳ Extracting: {filename} → {output_name}")
            cmd = [
                'ffmpeg', '-i', input_path,
                '-vn',
                '-c:a', 'libmp3lame' if output_format == 'mp3' else output_format,
                '-b:a', bitrate,
                '-y', output_path
            ]
            subprocess.run(cmd, capture_output=True)
            print(f"✓ Done: {output_name}")

# Usage
extract_audio_recursive("~/Videos/Recordings", output_format='mp3', bitrate='320k')
Parallel Processing
python
import os
import subprocess
from concurrent.futures import ThreadPoolExecutor, as_completed

def extract_audio_parallel(root_dir, workers=4):
    """Extract audio using multiple parallel workers."""
    video_files = []
    video_extensions = {'.mp4', '.mkv', '.webm'}
    
    for dirpath, _, filenames in os.walk(root_dir):
        for f in filenames:
            if os.path.splitext(f)[1].lower() in video_extensions:
                video_files.append(os.path.join(dirpath, f))
    
    def process_file(filepath):
        output = os.path.splitext(filepath)[0] + '.mp3'
        if os.path.exists(output):
            return f"✓ Skipped (exists): {os.path.basename(filepath)}"
        
        cmd = [
            'ffmpeg', '-i', filepath,
            '-vn', '-c:a', 'libmp3lame',
            '-b:a', '320k', '-y', output
        ]
        subprocess.run(cmd, capture_output=True, timeout=300)
        return f"✓ Extracted: {os.path.basename(filepath)}"
    
    with ThreadPoolExecutor(max_workers=workers) as executor:
        futures = {executor.submit(process_file, f): f for f in video_files}
        for future in as_completed(futures):
            print(future.result())

# Usage
extract_audio_parallel("~/Videos", workers=4)

Audio Normalization and Leveling

EBU R128 Loudness Standard
python
import subprocess
import json

def normalize_loudness(input_file, output_file, target_lufs=-16):
    """
    Normalize audio to target loudness using EBU R128 standard.
    
    Args:
        input_file: Source audio file
        output_file: Output file path
        target_lufs: Target loudness in LUFS (default: -16 for podcasts, -14 for music)
    """
    # First pass: measure loudness
    measure_cmd = [
        'ffmpeg', '-i', input_file,
        '-af', f'loudnorm=I={target_lufs}:LRA=11:TP=-1.5:print_format=json',
        '-f', 'null', '-'
    ]
    result = subprocess.run(measure_cmd, capture_output=True, text=True, timeout=60)
    
    # Second pass: apply normalization
    normalize_cmd = [
        'ffmpeg', '-i', input_file,
        '-af', f'loudnorm=I={target_lufs}:LRA=11:TP=-1.5',
        '-c:a', 'libmp3lame', '-b:a', '320k',
        '-y', output_file
    ]
    subprocess.run(normalize_cmd, capture_output=True, timeout=120)
    print(f"Normalized to {target_lufs} LUFS: {output_file}")

# Usage
normalize_loudness("input.mp3", "output.mp3", target_lufs=-16)

Splitting Audio by Chapters

Chapter-Based Splitting
python
import subprocess
import json

def split_by_chapters(input_file, output_dir="splits"):
    """
    Split an audio file into chapters using ffmpeg chapter metadata.
    """
    import os
    os.makedirs(output_dir, exist_ok=True)
    
    # Get chapter info
    cmd = [
        'ffprobe', '-i', input_file,
        '-print_format', 'json',
        '-show_chapters',
        '-loglevel', 'error'
    ]
    result = subprocess.run(cmd, capture_output=True, text=True)
    chapters = json.loads(result.stdout).get('chapters', [])
    
    if not chapters:
        print("No chapters found in the file.")
        return
    
    for chapter in chapters:
        start = chapter['start_time']
        end = chapter['end_time']
        title = chapter.get('tags', {}).get('title', f'Chapter {chapter["id"]}')
        
        safe_title = "".join(c for c in title if c.isalnum() or c in ' -_')
        output_path = os.path.join(output_dir, f"{safe_title}.mp3")
        
        cmd = [
            'ffmpeg', '-i', input_file,
            '-ss', str(start),
            '-to', str(end),
            '-c:a', 'libmp3lame', '-b:a', '320k',
            '-y', output_path
        ]
        subprocess.run(cmd, capture_output=True, timeout=300)
        print(f"✓ Split: {title} ({start}s → {end}s)")

# Usage
split_by_chapters("podcast.mp3", "~/Music/Splits")

Speech-to-Text Integration

Extracting Audio for Transcription
python
import subprocess
import os

def prepare_for_transcription(video_file, output_wav="speech.wav"):
    """
    Extract clean speech-optimized audio for transcription.
    Converts to mono 16kHz WAV (standard for speech recognition).
    """
    cmd = [
        'ffmpeg', '-i', video_file,
        '-vn',                    # No video
        '-acodec', 'pcm_s16le',   # 16-bit PCM
        '-ac', '1',               # Mono
        '-ar', '16000',           # 16kHz sample rate
        '-af', 'highpass=200,lowpass=8000',  # Speech frequency filter
        '-y', output_wav
    ]
    subprocess.run(cmd, capture_output=True, timeout=300)
    print(f"✓ Audio prepared for transcription: {output_wav}")
    return output_wav

# Usage
prepare_for_transcription("lecture.mp4", "lecture_audio.wav")

Skill Maturity Model

LevelCoverageQualityMetadataAutomation
1: BasicOne-off extractionsDefault qualityNoneManual
2: ConsistentFormat selection, basic batchTarget bitrateBasic tagsShell scripts
3: OrganizedBatch processing, normalizationOptimized per use caseFull ID3 + album artConfig presets
4: AutomatedWatch folders, scheduled jobsVerified qualityAutomatic taggingCron jobs + webhooks
5: LibraryFull pipeline, multi-format archiveLossless originals + playback copiesComplete metadata + coverFull automation with monitoring

Target: Level 3 for personal music collections. Level 4 for podcast production pipelines. Level 5 for media archiving at scale.


Show full SKILL.md (246 more words)Show less

Common Mistakes

  1. Transcoding lossy to lossy: Converting MP3 to FLAC doesn't restore quality — you get a large file with the same lossy audio. Always keep the original or use a lossless source.
  2. Ignoring sample rate and bit depth: For archival, use the source's native sample rate. Unnecessary resampling degrades quality. Only convert sample rates when needed.
  3. Missing album art in playable files: Many music players display album art prominently. Without it, your library looks unprofessional. Always embed cover art.
  4. Not normalizing volume levels: Different sources have vastly different loudness levels. Without normalization, switching between tracks or podcasts means constantly adjusting volume.
  5. Using wrong bitrate for the content: Music at 128kbps sounds noticeably compressed. Use 320kbps for music, 128kbps is fine for speech/podcasts. Opus at 96kbps is excellent for both.
  6. Overwriting originals: Always work on copies. A mistyped ffmpeg command can destroy your original file. Use -y cautiously.
  7. Neglecting metadata hygiene: Inconsistent or missing tags make a library unsearchable. Establish a tagging convention and stick to it.
  8. Not checking for clipping after normalization: Aggressive loudness normalization can cause clipping. Use true peak limiting (TP=-1.5 in loudnorm) to prevent this.
  9. Forgetting to test output quality: Don't trust the bitrate display — actually listen to a sample. Some extraction pipelines produce artifacts that aren't obvious in the metadata.
  10. Inconsistent naming schemes: A mix of conventions (Title.mp3 vs artist-title.mp3 vs track_number_title.mp3) makes automation harder. Pick one scheme and apply it universally.

© cosmicstack-labs, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in categories/media-download/audio-extraction of cosmicstack-labs/mercury-agent-skills.

Open the folder on GitHubat commit 30392fb

Compare with similar skills

Audio Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio Extraction this skillcosmicstack-labs/mercury-agent-skills476—~4.7kAutomated safety check: PassMIT
Gh Stackremotion-dev/remotion62k6 repos~2.4kAutomated safety check: PassCustom licence
HyperFrames Media Useheygen-com/hyperframes58k2 repos~2.1kAutomated safety check: PassApache-2.0
Guizang Social Cardsop7418/guizang-social-card-skill7.4k1 repos~7.8kAutomated safety check: PassAGPL-3.0
Weekly Changelog Videoheygen-com/hyperframes58k—~3.3kAutomated safety check: PassApache-2.0
Anthropic Brand Stylinganthropics/skills180k29 repos~559Automated safety check: PassApache-2.0

Similar skills

  • Gh Stack

    remotion-dev/remotion

    Official

    Manages stacked PRs and splits multi-part work into reviewable branches with gh-stack.

    62k GitHub starsUsed in 6 repos~2.4k tokens
    Media & CreativeAuto-check passed
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    58k GitHub starsUsed in 2 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • Guizang Social Cards

    op7418/guizang-social-card-skill

    Produces social card sets for Xiaohongshu and WeChat: carousels, Live Photo motion cards and puzzle layouts, and WeChat cover pairs, rendered from single-file HTML.

    7.4k GitHub starsUsed in 1 repo~7.8k tokens
    Media & CreativeAuto-check passed
  • Weekly Changelog Video

    heygen-com/hyperframes

    Turns a weekly changelog markdown file into a branded HyperFrames video with voiceover, animated mock-UI scenes and captions, using fonts, background and scripts bundled in the skill.

    58k GitHub stars~3.3k tokensUpdated today
    Media & CreativeAuto-check passed
  • Anthropic Brand Styling

    anthropics/skills

    Official

    Applies Anthropic's brand colors and fonts to artifacts such as PowerPoint slides, using fixed hex values for text and accents, Poppins headings and Lora body text.

    180k GitHub starsUsed in 29 repos~559 tokens
    Media & CreativeAuto-check passed
  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings

More from cosmicstack-labs/mercury-agent-skills

All 12 skills in this repo
  • Before You Build

    cosmicstack-labs/mercury-agent-skills

    Use this before implementing a product, feature, SaaS, AI app, or side project to score product risk and choose the smallest validation step.

    476 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Hyperframes CLI

    cosmicstack-labs/mercury-agent-skills

    HyperFrames CLI dev loop — project scaffolding, validation (lint/inspect), browser preview with live reload, MP4/WebM rendering, and environment troubleshooting (doctor, browser, info, upgrade).

    476 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Hyperframes Media

    cosmicstack-labs/mercury-agent-skills

    Asset preprocessing for HyperFrames compositions — local text-to-speech narration (Kokoro-82M, no API key), audio/video transcription (Whisper), and background removal for transparent overlays…

    476 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Handoff Protocols

    cosmicstack-labs/mercury-agent-skills

    Design and implement agent-to-agent handoff protocols for multi-agent systems.

    476 GitHub stars~4.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Health Monitoring

    cosmicstack-labs/mercury-agent-skills

    Monitor AI agent health, detect anomalies, set up alerting, and maintain observability dashboards for production multi-agent systems.

    476 GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Task Delegation

    cosmicstack-labs/mercury-agent-skills

    Design and operate task delegation systems for multi-agent fleets.

    476 GitHub stars~3.4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Audio Extraction

What does Audio Extraction do?

Audio Extraction: Extracting audio from videos, converting formats, and managing audio collections. Audio Extraction is an agent skill from cosmicstack-labs/mercury-agent-skills.

When should I use Audio Extraction?

Audio Extraction fits situations like: media & Creative work in your project.

How do I install Audio Extraction in Claude Code?

Run `npx skills add cosmicstack-labs/mercury-agent-skills --skill audio-extraction -a claude-code`. Or copy the skill folder (categories/media-download/audio-extraction in cosmicstack-labs/mercury-agent-skills) into .claude/skills/audio-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Audio Extraction in Codex?

Run `npx skills add cosmicstack-labs/mercury-agent-skills --skill audio-extraction -a codex`. Or copy the skill folder (categories/media-download/audio-extraction in cosmicstack-labs/mercury-agent-skills) into .agents/skills/audio-extraction in your project. Codex loads it when a task matches its description.

Can I use Audio Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add cosmicstack-labs/mercury-agent-skills --skill audio-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-extraction, .gemini/skills/audio-extraction, .github/skills/audio-extraction and .opencode/skills/audio-extraction in your project.

What does Audio Extraction need to run?

Going by SKILL.md and its folder, Audio Extraction needs the command-line tools its instructions call (yt-dlp, ffmpeg and pip). Our summary lists: Python 3.

Does Audio Extraction access the network?

SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Audio Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audio Extraction use?

Audio Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio Extraction use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio Extraction?

Skills that share tags, products or a category with Audio Extraction: Gh Stack (remotion-dev/remotion, 62k stars), HyperFrames Media Use (heygen-com/hyperframes, 58k stars), Guizang Social Cards (op7418/guizang-social-card-skill, 7.4k stars) and Weekly Changelog Video (heygen-com/hyperframes, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio Extraction?

cosmicstack-labs (a GitHub organization) maintains it in cosmicstack-labs/mercury-agent-skills, which has 476 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on August 25, 2026.

Source: cosmicstack-labs/mercury-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.