Agent skill

Audiobook

by benchflow-ai in benchflow-ai/skillsbench

Create audiobooks from web content or text files. An agent skill from benchflow-ai/skillsbench.

Apache-2.0Auto-check passedMedia & Creative

Install Audiobook

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill audiobook -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench audiobook --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks-extra/pg-essay-to-audiobook/environment/skills/audiobook .claude/skills/audiobook && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audiobook
GitHub stars
1.8k
Token cost
~2.4k tokens
SKILL.md length
230 words
Files
1
Skills in repo
178
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create audiobooks from web content or text files. An agent skill from benchflow-ai/skillsbench.

  • Works in 3 steps: Fetching Web Content → Text Processing → TTS Conversion with Fallback
  • Tasks that involve Text to speech and voice
  • SKILL.md covers Quick Start, Step 1: Fetching Web Content, Step 2: Text Processing and Step 3: TTS Conversion with…, plus 3 more sections
  • Reaches api.elevenlabs.io and api.openai.com; needs OPENAI_API_KEY and ELEVENLABS_API_KEY

What it does

Audiobook is an agent skill from benchflow-ai/skillsbench. Create audiobooks from web content or text files. Handles content fetching, text processing, and TTS conversion with automatic fallback between ElevenLabs, OpenAI TTS, and gTTS.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with ElevenLabs and OpenAI. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “/audiobook”

Requirements

  • Python 3
  • A credential in ELEVENLABS_API_KEY
  • A credential in OPENAI_API_KEY

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Fetching Web Content
  2. Text Processing
  3. TTS Conversion with Fallback

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.elevenlabs.io
    • api.openai.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ELEVENLABS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audiobook loads about 2.4k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 230 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 230 words, ~2,402 tokens.

Download SKILL.mdSave it as .claude/skills/audiobook/SKILL.md (or your agent's skills folder).
name
audiobook
description
Create audiobooks from web content or text files. Handles content fetching, text processing, and TTS conversion with automatic fallback between ElevenLabs, OpenAI TTS, and gTTS.

Audiobook Creation Guide

Create audiobooks from web articles, essays, or text files. This skill covers the full pipeline: content fetching, text processing, and audio generation.

Quick Start

python
import os

# 1. Check which TTS API is available
def get_tts_provider():
    if os.environ.get("ELEVENLABS_API_KEY"):
        return "elevenlabs"
    elif os.environ.get("OPENAI_API_KEY"):
        return "openai"
    else:
        return "gtts"  # Free, no API key needed

provider = get_tts_provider()
print(f"Using TTS provider: {provider}")

Step 1: Fetching Web Content

IMPORTANT: Verify fetched content is complete

WebFetch and similar tools may return summaries instead of full text. Always verify:

python
import subprocess

def fetch_article_content(url):
    """Fetch article content using curl for reliability."""
    # Use curl to get raw HTML - more reliable than web fetch tools
    result = subprocess.run(
        ["curl", "-s", url],
        capture_output=True,
        text=True
    )
    html = result.stdout

    # Strip HTML tags (basic approach)
    import re
    text = re.sub(r'<script[^>]*>.*?</script>', '', html, flags=re.DOTALL)
    text = re.sub(r'<style[^>]*>.*?</style>', '', html, flags=re.DOTALL)
    text = re.sub(r'<[^>]+>', ' ', text)
    text = re.sub(r'\s+', ' ', text).strip()

    return text
Content verification checklist

Before converting to audio, verify:

  • Text length is reasonable for the source (articles typically 1,000-10,000+ words)
  • Content includes actual article text, not just navigation/headers
  • No "summary" or "key points" headers that indicate truncation
python
def verify_content(text, expected_min_chars=1000):
    """Basic verification that content is complete."""
    if len(text) < expected_min_chars:
        print(f"WARNING: Content may be truncated ({len(text)} chars)")
        return False
    if "summary" in text.lower()[:500] or "key points" in text.lower()[:500]:
        print("WARNING: Content appears to be a summary, not full text")
        return False
    return True

Step 2: Text Processing

Clean and prepare text for TTS
python
import re

def clean_text_for_tts(text):
    """Clean text for better TTS output."""
    # Remove URLs
    text = re.sub(r'http[s]?://\S+', '', text)

    # Remove footnote markers like [1], [2]
    text = re.sub(r'\[\d+\]', '', text)

    # Normalize whitespace
    text = re.sub(r'\s+', ' ', text)

    # Remove special characters that confuse TTS
    text = re.sub(r'[^\w\s.,!?;:\'"()-]', '', text)

    return text.strip()

def chunk_text(text, max_chars=4000):
    """Split text into chunks at sentence boundaries."""
    sentences = re.split(r'(?<=[.!?])\s+', text)
    chunks = []
    current_chunk = ""

    for sentence in sentences:
        if len(current_chunk) + len(sentence) < max_chars:
            current_chunk += sentence + " "
        else:
            if current_chunk:
                chunks.append(current_chunk.strip())
            current_chunk = sentence + " "

    if current_chunk:
        chunks.append(current_chunk.strip())

    return chunks

Step 3: TTS Conversion with Fallback

Automatic provider selection
python
import os
import subprocess

def create_audiobook(text, output_path):
    """Convert text to audiobook with automatic TTS provider selection."""

    # Check available providers
    has_elevenlabs = bool(os.environ.get("ELEVENLABS_API_KEY"))
    has_openai = bool(os.environ.get("OPENAI_API_KEY"))

    if has_elevenlabs:
        print("Using ElevenLabs TTS (highest quality)")
        return create_with_elevenlabs(text, output_path)
    elif has_openai:
        print("Using OpenAI TTS (high quality)")
        return create_with_openai(text, output_path)
    else:
        print("Using gTTS (free, no API key required)")
        return create_with_gtts(text, output_path)
ElevenLabs implementation
python
import requests

def create_with_elevenlabs(text, output_path):
    """Generate audiobook using ElevenLabs API."""
    api_key = os.environ.get("ELEVENLABS_API_KEY")
    voice_id = "21m00Tcm4TlvDq8ikWAM"  # Rachel - calm female voice

    chunks = chunk_text(text, max_chars=4500)
    audio_files = []

    for i, chunk in enumerate(chunks):
        chunk_file = f"/tmp/chunk_{i:03d}.mp3"

        response = requests.post(
            f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}",
            headers={
                "xi-api-key": api_key,
                "Content-Type": "application/json"
            },
            json={
                "text": chunk,
                "model_id": "eleven_turbo_v2_5",
                "voice_settings": {"stability": 0.5, "similarity_boost": 0.75}
            }
        )

        if response.status_code == 200:
            with open(chunk_file, "wb") as f:
                f.write(response.content)
            audio_files.append(chunk_file)
        else:
            print(f"Error: {response.status_code} - {response.text}")
            return False

    return concatenate_audio(audio_files, output_path)
OpenAI TTS implementation
python
def create_with_openai(text, output_path):
    """Generate audiobook using OpenAI TTS API."""
    api_key = os.environ.get("OPENAI_API_KEY")

    chunks = chunk_text(text, max_chars=4000)
    audio_files = []

    for i, chunk in enumerate(chunks):
        chunk_file = f"/tmp/chunk_{i:03d}.mp3"

        response = requests.post(
            "https://api.openai.com/v1/audio/speech",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json"
            },
            json={
                "model": "tts-1",
                "input": chunk,
                "voice": "onyx",  # Deep male voice, good for essays
                "response_format": "mp3"
            }
        )

        if response.status_code == 200:
            with open(chunk_file, "wb") as f:
                f.write(response.content)
            audio_files.append(chunk_file)
        else:
            print(f"Error: {response.status_code} - {response.text}")
            return False

    return concatenate_audio(audio_files, output_path)
gTTS implementation (free fallback)
python
def create_with_gtts(text, output_path):
    """Generate audiobook using gTTS (free, no API key)."""
    from gtts import gTTS
    from pydub import AudioSegment

    chunks = chunk_text(text, max_chars=4500)
    audio_files = []

    for i, chunk in enumerate(chunks):
        chunk_file = f"/tmp/chunk_{i:03d}.mp3"

        tts = gTTS(text=chunk, lang='en', slow=False)
        tts.save(chunk_file)
        audio_files.append(chunk_file)

    return concatenate_audio(audio_files, output_path)
Audio concatenation
python
def concatenate_audio(audio_files, output_path):
    """Concatenate multiple audio files using ffmpeg."""
    if not audio_files:
        return False

    # Create file list for ffmpeg
    list_file = "/tmp/audio_list.txt"
    with open(list_file, "w") as f:
        for audio_file in audio_files:
            f.write(f"file '{audio_file}'\n")

    # Concatenate with ffmpeg
    result = subprocess.run([
        "ffmpeg", "-y", "-f", "concat", "-safe", "0",
        "-i", list_file, "-c", "copy", output_path
    ], capture_output=True)

    # Cleanup temp files
    import os
    for f in audio_files:
        os.unlink(f)
    os.unlink(list_file)

    return result.returncode == 0

Complete Example

python
#!/usr/bin/env python3
"""Create audiobook from web articles."""

import os
import re
import subprocess
import requests

# ... include all helper functions above ...

def main():
    # Fetch articles
    urls = [
        "https://example.com/article1",
        "https://example.com/article2"
    ]

    all_text = ""
    for url in urls:
        print(f"Fetching: {url}")
        text = fetch_article_content(url)

        if not verify_content(text):
            print(f"WARNING: Content from {url} may be incomplete")

        all_text += f"\n\n{text}"

    # Clean and convert
    clean_text = clean_text_for_tts(all_text)
    print(f"Total text: {len(clean_text)} characters")

    # Create audiobook
    success = create_audiobook(clean_text, "/root/audiobook.mp3")

    if success:
        print("Audiobook created successfully!")
    else:
        print("Failed to create audiobook")

if __name__ == "__main__":
    main()

TTS Provider Comparison

ProviderQualityCostAPI Key RequiredBest For
ElevenLabsExcellentPaidYesProfessional audiobooks
OpenAI TTSVery GoodPaidYesGeneral purpose
gTTSGoodFreeNoTesting, budget projects

Troubleshooting

"Content appears to be a summary"
  • Use curl directly instead of web fetch tools
  • Verify the URL is correct and accessible
  • Check if the site requires JavaScript rendering
"API key not found"
  • Check environment variables: echo $OPENAI_API_KEY
  • Ensure keys are exported in the shell
  • Fall back to gTTS if no paid API keys available
"Audio chunks don't sound continuous"
  • Ensure chunking happens at sentence boundaries
  • Consider adding small pauses between sections
  • Use consistent voice settings across all chunks

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks-extra/pg-essay-to-audiobook/environment/skills/audiobook of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Audiobook next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audiobook compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audiobook this skillbenchflow-ai/skillsbench1.8k—~2.4kAutomated safety check: PassApache-2.0
Video Translatorshang-zhu/violin1.1k—~1kAutomated safety check: NotesMIT
9Router Text to Speechdecolua/9router30k—~765Automated safety check: PassMIT
Super Video MakerBomx/super-video-maker-skill308—~11kAutomated safety check: NotesNone
Motion Videobestagentkits/motion-video-skill113—~1.5kAutomated safety check: PassMIT
Local AI Useamd/skills398—~5kAutomated safety check: NotesMIT

Similar skills

  • Video Translator

    shang-zhu/violin

    Dub a video into another language and generate subtitles using the default Together + Cartesia stack.

    1.1k GitHub stars~1k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • 9Router Text to Speech

    decolua/9router

    Turns text into spoken audio through a 9Router server, choosing among voices from OpenAI, ElevenLabs, Deepgram, Edge TTS and other providers.

    30k GitHub stars~765 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Super Video Maker

    Bomx/super-video-maker-skill

    End-to-end AI video production skill for agentic frameworks.

    308 GitHub stars~11k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Motion Video

    bestagentkits/motion-video-skill

    Produce beat-synced 1080p motion-graphic videos in HyperFrames (HTML + GSAP) with an AI voice-over, Vietnamese karaoke captions, SFX and generated music, in one of two proven styles (glass keynote…

    113 GitHub stars~1.5k tokensUpdated 10 days ago
    Media & CreativeAuto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Voice AI Development

    davila7/claude-code-templates

    Expert in building voice AI applications - from real-time voice agents to voice-enabled apps.

    32k GitHub starsUsed in 6 repos~2.1k tokens
    Media & CreativeAuto-check passed

More from benchflow-ai/skillsbench

All 178 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Audiobook

What does Audiobook do?

Create audiobooks from web content or text files. An agent skill from benchflow-ai/skillsbench. Audiobook is an agent skill from benchflow-ai/skillsbench. Create audiobooks from web content or text files.

When should I use Audiobook?

Audiobook fits situations like: tasks that involve Text to speech and voice.

How do I install Audiobook in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill audiobook -a claude-code`. Or copy the skill folder (tasks-extra/pg-essay-to-audiobook/environment/skills/audiobook in benchflow-ai/skillsbench) into .claude/skills/audiobook in your project. Claude Code loads it when a task matches its description.

How do I install Audiobook in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill audiobook -a codex`. Or copy the skill folder (tasks-extra/pg-essay-to-audiobook/environment/skills/audiobook in benchflow-ai/skillsbench) into .agents/skills/audiobook in your project. Codex loads it when a task matches its description.

Can I use Audiobook in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill audiobook -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audiobook, .gemini/skills/audiobook, .github/skills/audiobook and .opencode/skills/audiobook in your project.

What does Audiobook need to run?

Going by SKILL.md and its folder, Audiobook needs credentials named OPENAI_API_KEY and ELEVENLABS_API_KEY. Our summary lists: Python 3; A credential in ELEVENLABS_API_KEY; A credential in OPENAI_API_KEY.

Does Audiobook access the network?

SKILL.md names 2 domains. In commands or code: api.elevenlabs.io and api.openai.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Audiobook safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Audiobook use?

Audiobook is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audiobook use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audiobook?

Skills that share tags, products or a category with Audiobook: Video Translator (shang-zhu/violin, 1.1k stars), 9Router Text to Speech (decolua/9router, 30k stars), Super Video Maker (Bomx/super-video-maker-skill, 308 stars) and Motion Video (bestagentkits/motion-video-skill, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audiobook?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 178 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.