Agent skill

Blog Audio

by AgriciDaniel in AgriciDaniel/claude-blog

Generate audio narration of blog posts using Google Gemini TTS.

MITAuto-check: notesMedia & Creative

Install Blog Audio

skills CLI
$ npx skills add AgriciDaniel/claude-blog --skill blog-audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AgriciDaniel/claude-blog blog-audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AgriciDaniel/claude-blog.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/blog-audio .claude/skills/blog-audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
blog-audio
GitHub stars
2.3k
Used in
1 other repo
Token cost
~2.2k tokens
SKILL.md length
895 words
Files
8 (incl. scripts, references)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Generate audio narration of blog posts using Google Gemini TTS.

  • Works in 6 steps: Read the Blog Post → Choose Mode → Prepare Text → …
  • User says blog audio
  • SKILL.md covers Quick Reference, Prerequisites, Always Use run.py Wrapper and API Key Check (Gate Pattern), plus 7 more sections
  • Runs Python scripts from its folder; calls python3 and apt; needs GOOGLE_AI_API_KEY

What it does

Blog Audio is an agent skill from AgriciDaniel/claude-blog. Generate audio narration of blog posts using Google Gemini TTS. Supports summary narration, full article read-aloud, and two-speaker podcast/dialogue mode with 30 voice options. Outputs MP3 with HTML5 audio embed code. Works standalone via /blog audio or internally from blog-write. Falls back gracefully when API key is not configured. Use when user says "blog audio", "narrate blog", "audio version", "text to speech", "tts", "podcast mode", "read aloud", "audio narration", "voice", "narration", "generate audio".

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/voices.md`, `scripts/__init__.py` and `scripts/generate_audio.py`).

It sits in Media & Creative, covering Text to speech and voice and Blog and article writing. It works with Google Gemini. The repository describes itself as: Claude Code blog skill suite: 30 sub-skills, 5 agents, 5-gate v1.9.0 Blog Delivery Contract, dual-optimized for Google rankings and AI citations. Active development at… The licence is MIT.

When your agent uses it

  • User says blog audio
  • Audio narration

Example prompts

  • “blog audio”
  • “narrate blog”
  • “audio version”
  • “/blog-audio”

Requirements

  • Python 3
  • A credential in GOOGLE_AI_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Read the Blog Post
  2. Choose Mode
  3. Prepare Text
  4. Select Voice
  5. Generate Audio
  6. Deliver

What it can do on your machine

Read from SKILL.md and the folder at commit 2500d4c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 6 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • apt

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GOOGLE_AI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Blog Audio loads about 2.2k tokens when it runs, and up to ~3.5k if it reads all its reference files. Until then it costs about 132 tokens; SKILL.md has 895 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:241
    | FFmpeg not found | Install: `sudo apt install ffmpeg`. Falls back to WAV output. |

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from AgriciDaniel/claude-blog at commit 2500d4c, republished under its MIT licence (© AgriciDaniel). 895 words, ~2,198 tokens.

Download SKILL.mdSave it as .claude/skills/blog-audio/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
blog-audio
description
Generate audio narration of blog posts using Google Gemini TTS. Supports summary narration, full article read-aloud, and two-speaker podcast/dialogue mode with 30 voice options. Outputs MP3 with HTML5 audio embed code. Works standalone via /blog audio or internally from blog-write. Falls back gracefully when API key is not configured. Use when user says "blog audio", "narrate blog", "audio version", "text to speech", "tts", "podcast mode", "read aloud", "audio narration", "voice", "narration", "generate audio".
user-invokable
true
argument-hint
[generate|voices|setup] [file-or-text] [--mode summary|full|dialogue] [--voice name]
license
MIT
metadata.author
AgriciDaniel
metadata.version
2.2.0

Blog Audio: Gemini TTS Narration for Blog Posts

Generate professional audio narration of blog content using Google's Gemini TTS. Three modes: summary (200-300 word spoken overview), full article read-aloud, or two-speaker podcast dialogue. 30 voices, 80+ languages, HTML5 embed output.

Quick Reference

CommandWhat it does
/blog audio generate <file>Generate audio narration of a blog post
/blog audio voicesShow available voices with characteristics
/blog audio setupCheck/configure API key for Gemini TTS

Prerequisites

  • Python 3.11+ (venv managed automatically by run.py)
  • GOOGLE_AI_API_KEY environment variable (same key used by blog-image)
  • FFmpeg (for WAV-to-MP3 conversion; falls back to WAV if missing)

Always Use run.py Wrapper

bash
# CORRECT:
python3 scripts/run.py generate_audio.py --text "..." --voice Charon --json

# WRONG:
python3 scripts/generate_audio.py --text "..."  # Fails without venv

API Key Check (Gate Pattern)

Before generating audio, check for the API key:

bash
test -n "${GOOGLE_AI_API_KEY:-}" && echo "GOOGLE_AI_API_KEY is set" || echo "GOOGLE_AI_API_KEY is not set"
  • If set: proceed with generation
  • If not set: guide the user: "Audio generation requires a Google AI API key. Get one free at https://aistudio.google.com/apikey Then set it: export GOOGLE_AI_API_KEY=your-key This can be the same key used by /blog image, but it must be exported in the shell."
  • When called internally (from blog-write): return silently if key is missing. Never block the writing workflow.

Setup

For /blog audio setup:

  1. Check if GOOGLE_AI_API_KEY is set in environment
  2. If blog-image uses project .mcp.json, confirm the referenced env var is exported
  3. If not, guide user to https://aistudio.google.com/apikey
  4. Verify with a dry run: python3 scripts/run.py generate_audio.py --text "Test" --dry-run --json

Voice Selection

For /blog audio voices:

Load references/voices.md and present the voice catalog to the user.

Ask the user which voice they prefer, or recommend based on content type:

  • Article narration: Charon (Informative) or Sadaltager (Knowledgeable)
  • Tutorial/how-to: Achird (Friendly) or Sulafat (Warm)
  • News/analysis: Rasalgethi (Informative) or Schedar (Even)
  • Lifestyle/wellness: Aoede (Breezy) or Vindemiatrix (Gentle)
  • Dialogue host: Puck (Upbeat) or Laomedeia (Upbeat)
  • Dialogue expert: Kore (Firm) or Charon (Informative)

Generation Workflow

For /blog audio generate <file>:

Step 1: Read the Blog Post

Read the file and extract:

  • Title (from H1 or frontmatter)
  • Full content (markdown body)
  • Approximate word count
Step 2: Choose Mode

Ask the user (or auto-select if they specified --mode):

ModeWhen to useOutput
SummaryQuick audio overview (1-2 min)200-300 word spoken summary
FullComplete read-aloud (5-15 min)Full article as natural speech
DialoguePodcast-style (3-8 min)Two-person conversation about the article
Step 3: Prepare Text

Claude prepares the text; the script does TTS only.

Summary mode: Write a 200-300 word spoken summary of the article. Rules:

  • Write as natural speech, not written text
  • Open with the article's key finding or answer
  • Cover 3-5 main takeaways
  • Close with actionable advice
  • No markdown, no "In this article...", no meta-commentary
  • Use conversational transitions ("Here's what matters...", "The key finding is...")

Full mode: Strip the markdown content to clean spoken text:

  • Headings become natural transitions ("Next, let's look at...")
  • Links become plain text (remove URLs, keep anchor text)
  • Images and charts: omit or briefly describe ("As the data shows...")
  • Code blocks: describe verbally ("The code uses a for-loop to...")
  • Lists: convert to natural sentences
  • Remove frontmatter, schema markup, HTML tags
  • Add brief intro: "This is [title], published on [date]."

Dialogue mode: Write a 2-person conversation script about the article:

  • Speaker1 = Host (curious, asks good questions)
  • Speaker2 = Expert (knowledgeable, gives clear answers)
  • Format each line as: Speaker1: What's the key takeaway here?
  • Cover the article's main points conversationally
  • 15-25 exchanges (produces ~3-8 minutes)
  • Natural, not stilted ("That's a great point" over "Indeed, as the research indicates")
Show full SKILL.md (334 more words)Show less
Step 4: Select Voice

If the user chose a voice, use it. Otherwise, recommend based on mode:

  • Summary/Full: default to Charon (Informative)
  • Dialogue: default to Puck (Host) + Kore (Expert)
Step 5: Generate Audio

Write the prepared text to a file under the working directory, then call:

bash
# Single voice (summary or full mode)
python3 scripts/run.py generate_audio.py \
  --text-file blog_audio_prepared.txt \
  --voice Charon \
  --model flash \
  --output audio/post-slug.mp3 \
  --json

# Two voices (dialogue mode)
python3 scripts/run.py generate_audio.py \
  --text-file blog_audio_dialogue.txt \
  --voice Puck \
  --voice2 Kore \
  --model pro \
  --output audio/post-slug-dialogue.mp3 \
  --json

Model selection:

  • flash (default): maps to gemini-3.1-flash-tts-preview, good for summaries and standard narration.
  • flash31: explicit alias for gemini-3.1-flash-tts-preview.
  • legacy-flash25: retained only for older compatibility.
  • pro or legacy-pro25: maps to gemini-2.5-pro-preview-tts, use only when needed.
Step 6: Deliver

Present the result to the user:

  1. File path: where the audio was saved
  2. Duration: human-readable (e.g., "3:42")
  3. Embed code: ready-to-paste HTML5 audio tag
  4. Cost: estimated API cost
  5. Placement suggestion: where to insert the embed in the blog post

Embedding Guide

Standard HTML (Hugo, Jekyll, static sites)
html
<audio controls preload="metadata">
  <source src="audio/post-slug.mp3" type="audio/mpeg">
  Your browser does not support the audio element.
</audio>
MDX (Next.js, Gatsby)
jsx
<audio controls preload="metadata">
  <source src="/audio/post-slug.mp3" type="audio/mpeg" />
</audio>
WordPress
[audio src="audio/post-slug.mp3"]
Placement

Insert the audio player after the introduction (below the first H2) or at the very top of the article with a label: "Listen to this article" or "Audio version".

Internal API (for blog-write)

When invoked internally from blog-write:

Input:

  • text: Prepared text (already cleaned by Claude)
  • voice: Voice name (default: Charon)
  • voice2: Second voice for dialogue (optional)
  • model: flash or pro
  • output_path: Where to save the file

Output:

markdown
### Audio Narration
- **Path:** /path/to/audio/post-slug.mp3
- **Duration:** 3:42
- **Voice:** Charon
- **Embed:** `<audio controls preload="metadata"><source src="audio/post-slug.mp3" type="audio/mpeg"></audio>`

Graceful fallback: If GOOGLE_AI_API_KEY is not set, return immediately with no error. The writing workflow continues without audio. Never block blog-write because audio generation is unavailable.

Error Handling

ErrorResolution
GOOGLE_AI_API_KEY not setGet key at https://aistudio.google.com/apikey
FFmpeg not foundInstall: sudo apt install ffmpeg. Falls back to WAV output.
Rate limitedWait and retry. Check limits at https://aistudio.google.com/rate-limit
Text too long (>8,192 input tokens)Split into sections around 7,800 tokens; the script chunks and stitches prepared text
Unknown voice nameRun /blog audio voices to see valid options
API errorCheck key validity and model availability
API key missing (internal call)Return silently: writing workflow continues

Reference Documentation

Load on-demand: do NOT load all at startup:

  • references/voices.md: Full 30-voice catalog, recommendations by content type, dialogue pairings

© AgriciDaniel, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in skills/blog-audio of AgriciDaniel/claude-blog.

  • SKILL.md
  • references/voices.md
  • scripts/__init__.py
  • scripts/generate_audio.py
  • scripts/requirements.lock
  • scripts/requirements.txt
  • scripts/run.py
  • scripts/setup_environment.py

Open the folder on GitHubat commit 2500d4c

Used in 1 other repository

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in AgriciDaniel/claude-blog, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Blog Audio next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Blog Audio compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Blog Audio this skillAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Short Video Scripteraaron-he-zhu/aaron-marketing-skills2.9k—~3.5kAutomated safety check: PassApache-2.0
Gemini Ttsiurysza/module-graph420—~968Automated safety check: PassMIT
Gemini Audioeinverne/dotfiles121—~2kAutomated safety check: NotesMIT
WeChat Article Publisherjiji262/wechat-publisher274—~4.4kAutomated safety check: PassNone
Video Analyzermikefutia/claude-vision101—~747Automated safety check: NotesNone

Similar skills

  • Short Video Scripter

    aaron-he-zhu/aaron-marketing-skills

    A skill your agent uses when the user asks to "script this short video", "write a TikTok / Reels / Shorts script", "给这条抖音或视频号视频写脚本", or "fix the hook — viewers drop off in the first seconds"…

    2.9k GitHub stars~3.5k tokensUpdated today
    Media & CreativeAuto-check passed
  • Gemini Tts

    iurysza/module-graph

    Generates spoken MP3 audio from text or Markdown with Gemini TTS.

    420 GitHub stars~968 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Gemini Audio

    einverne/dotfiles

    Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

    121 GitHub stars~2k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • WeChat Article Publisher

    jiji262/wechat-publisher

    Researches a topic, writes an illustrated WeChat Official Account article in a chosen author voice and sends it to the account's draft box.

    274 GitHub stars~4.4k tokensUpdated 2 mo ago
    Writing & ContentAuto-check passed
  • Video Analyzer

    mikefutia/claude-vision

    Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details…

    101 GitHub stars~747 tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Fal AI

    mikeOnBreeze/cc-crossbeam

    This skill enables AI video generation from images AND text-to-speech voiceover generation using Fal.ai's API.

    293 GitHub stars~1.9k tokensUpdated 7 mo ago
    Media & CreativeAuto-check: notes

More from AgriciDaniel/claude-blog

All 39 skills in this repo
  • Blog Google

    AgriciDaniel/claude-blog

    Google API integration for blog performance: PageSpeed Insights, CrUX Core Web Vitals with 25-week history, Search Console performance, URL Inspection, Indexing API, GA4 organic traffic, NLP entity…

    2.3k GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check: notes
  • Blog Flow

    AgriciDaniel/claude-blog

    FLOW framework integration for bloggers. An agent skill from AgriciDaniel/claude-blog.

    2.3k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Blog Notebooklm

    AgriciDaniel/claude-blog

    Query Google NotebookLM notebooks for source-grounded, citation-backed answers from user-uploaded documents.

    2.3k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check: warnings
  • Blog Image

    AgriciDaniel/claude-blog

    AI image generation and editing for blog content powered by Gemini via MCP.

    2.3k GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Blog Cluster

    AgriciDaniel/claude-blog

    Semantic topic cluster planning and automated execution engine for claude-blog.

    2.3k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Blog Discourse

    AgriciDaniel/claude-blog

    Research what people are actually saying about a topic in the last 30 days across Reddit, X / Twitter, YouTube, Hacker News, dev.to, Medium, and other public discourse platforms.

    2.3k GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check: warnings

Works with

Questions about Blog Audio

What does Blog Audio do?

Generate audio narration of blog posts using Google Gemini TTS. Blog Audio is an agent skill from AgriciDaniel/claude-blog. Generate audio narration of blog posts using Google Gemini TTS.

When should I use Blog Audio?

Blog Audio fits situations like: user says blog audio; audio narration.

How do I install Blog Audio in Claude Code?

Run `npx skills add AgriciDaniel/claude-blog --skill blog-audio -a claude-code`. Or copy the skill folder (skills/blog-audio in AgriciDaniel/claude-blog) into .claude/skills/blog-audio in your project. Claude Code loads it when a task matches its description.

How do I install Blog Audio in Codex?

Run `npx skills add AgriciDaniel/claude-blog --skill blog-audio -a codex`. Or copy the skill folder (skills/blog-audio in AgriciDaniel/claude-blog) into .agents/skills/blog-audio in your project. Codex loads it when a task matches its description.

Can I use Blog Audio in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AgriciDaniel/claude-blog --skill blog-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/blog-audio, .gemini/skills/blog-audio, .github/skills/blog-audio and .opencode/skills/blog-audio in your project.

What does Blog Audio need to run?

Going by SKILL.md and its folder, Blog Audio needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and apt) and credentials named GOOGLE_AI_API_KEY. Our summary lists: Python 3; A credential in GOOGLE_AI_API_KEY.

Does Blog Audio access the network?

SKILL.md names 1 domain. As links in the text: aistudio.google.com. This is read from the text; nothing was executed.

Is Blog Audio safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Blog Audio use?

Blog Audio is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Blog Audio use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Blog Audio?

Skills that share tags, products or a category with Blog Audio: Short Video Scripter (aaron-he-zhu/aaron-marketing-skills, 2.9k stars), Gemini Tts (iurysza/module-graph, 420 stars), Gemini Audio (einverne/dotfiles, 121 stars) and WeChat Article Publisher (jiji262/wechat-publisher, 274 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Blog Audio?

AgriciDaniel (a GitHub user) maintains it in AgriciDaniel/claude-blog, which has 2,348 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on October 9, 2026.

Source: AgriciDaniel/claude-blog on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.