Agent skill

Story Narrator

by hassancs91 in hassancs91/claude-image-generation

Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech).

MITAuto-check passedMedia & Creative

Install Story Narrator

skills CLI
$ npx skills add hassancs91/claude-image-generation --skill story-narrator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hassancs91/claude-image-generation story-narrator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hassancs91/claude-image-generation.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/story-narrator .claude/skills/story-narrator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
story-narrator
GitHub stars
102
Token cost
~2.5k tokens
SKILL.md length
1,249 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech).

  • Works in 6 steps: Load scenes.json → Intake and analysis → Voice strategy and selection → …
  • The user wants to narrate a story
  • SKILL.md covers When this skill applies, Required tool — the ElevenLabs…, Workflow — five stages with… and Cost expectations, plus 2 more sections
  • Calls uvx; needs ELEVENLABS_API_KEY

What it does

Story Narrator is an agent skill from hassancs91/claude-image-generation. Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech). Reads {slug}scenes.json (from scene-splitter), proposes a warm storyteller voice, drafts the narration text per scene (optionally with Eleven v3 audio tags for emotion), lets the user review, then generates one MP3 per scene saved as {slug}partNN.mp3 so it pairs by index with the scene's image. Use this skill whenever the user wants to narrate a story, generate per-scene audio, voice a storybook…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. It works with ElevenLabs, Model Context Protocol and Storybook. The repository describes itself as: Connect Claude to image generation with Agent Skills. Three levels: a zero-cost code-based design engine, a Three.js 3D renderer, and a real diffusion model on Cloudflare. Plus… The licence is MIT.

When your agent uses it

  • The user wants to narrate a story
  • Generate per-scene audio
  • Voice a storybook
  • Create read-along narration

Example prompts

  • “narrate this story”
  • “generate the audio”
  • “voice the storybook”
  • “/story-narrator”

Requirements

  • A credential in ELEVENLABS_API_KEY

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Load scenes.json
  2. Intake and analysis
  3. Voice strategy and selection
  4. Narration script (+ optional emotion tags)
  5. HARD GATE: review the narration script
  6. Generate audio (one MP3 per scene)

What it can do on your machine

Read from SKILL.md and the folder at commit f533831. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uvx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ELEVENLABS_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Story Narrator loads about 2.5k tokens when it runs. Until then it costs about 212 tokens; SKILL.md has 1,249 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~212
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hassancs91/claude-image-generation at commit f533831, republished under its MIT licence (© hassancs91). 1,249 words, ~2,501 tokens.

Download SKILL.mdSave it as .claude/skills/story-narrator/SKILL.md (or your agent's skills folder).
name
story-narrator
description
Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (text_to_speech). Reads {slug}_scenes.json (from scene-splitter), proposes a warm storyteller voice, drafts the narration text per scene (optionally with Eleven v3 audio tags for emotion), lets the user review, then generates one MP3 per scene saved as {slug}_part_NN.mp3 so it pairs by index with the scene's image. Use this skill whenever the user wants to narrate a story, generate per-scene audio, voice a storybook, create read-along narration, or produce TTS for the storybook pipeline. Trigger on phrases like "narrate this story", "generate the audio", "voice the storybook", "make the narration", "read this story aloud", or whenever scenes.json exists and the user wants spoken audio. Requires the ElevenLabs MCP connected.

Story Narrator

Turns a story (already split into scenes) into a set of expressive narration clips — one MP3 per scene — using the ElevenLabs MCP. Each clip is saved as {slug}_part_NN.mp3 so it pairs by index with the scene's image: scene 3's picture and scene 3's narration play together in the final storybook.

When this skill applies

{slug}_scenes.json exists (from scene-splitter) and the user wants per-scene audio. Also applies if the user pastes a story and asks to narrate it — split it first (or narrate paragraph-by-paragraph).

The skill does NOT apply to:

  • A single short utterance (call text_to_speech directly).
  • Real-time conversational TTS, music, or sound effects.
  • Voice cloning (do that separately in ElevenLabs, then bring the voice_id here).

Required tool — the ElevenLabs MCP

The skill calls mcp__elevenlabs__text_to_speech. Confirm it's available before Stage 4. If the MCP isn't connected, tell the user to install it (uvx elevenlabs-mcp with ELEVENLABS_API_KEY set — see https://github.com/elevenlabs/elevenlabs-mcp) and stop. Stages 1–3 (voice choice + narration script) work without it; only generation needs it.

Key text_to_speech parameters this skill uses:

  • text — the scene's narration text
  • voice_id (or voice_name) — the chosen storyteller voice
  • model_id — see Stage 2 (default a v3 model for audio-tag expressiveness; fall back to eleven_multilingual_v2)
  • stability — 0.4–0.5 for natural, expressive delivery (lower = more emotional range)
  • style — small positive value (e.g. 0.2) adds expressiveness; 0 is flat
  • output_directory — set to the story's audio folder so files land in the right place

The MCP saves the file and returns its path — it names the file itself, so this skill renames each result to the locked {slug}_part_NN.mp3 convention after generation (Stage 5).

Workflow — five stages with one hard gate

Stage 0: Load scenes.json

Read {slug}_scenes.json. Treat each scene's text as one narration clip. Use each scene's mood to guide delivery/tag choices. The clip for scenes[i].index = N becomes {slug}_part_NN.mp3.

Stage 1: Intake and analysis
  1. Read all scene texts.
  2. Note total characters (rough cost = chars × ~$0.10/1k) and how much is dialogue.
  3. One-line summary: "N scenes, ~C total characters, ~D% dialogue. Recommendation: single narrator."
Stage 2: Voice strategy and selection

Propose single narrator (the default and best choice for beginner storybooks — one warm voice carrying the whole story, shifting tone for dialogue). Only consider per-character voices if dialogue is heavy AND there are 2+ distinct recurring speakers AND their voices should clearly differ — and even then, single narrator usually sounds more cohesive for a short children's story.

Voice choice:

  • If the user already has a preferred voice_id, use it.
  • Otherwise, suggest finding a warm, friendly storyteller voice. You can call mcp__elevenlabs__search_voices (e.g. search "storyteller" or "warm narration") and propose ONE specific voice with its voice_id. Don't list five.
  • Ask once if needed: "What voice_id should I use? (Find one in your ElevenLabs library, or I can search for a warm storyteller voice.)" Suggest the user save it for future runs.

Model choice:

  • Default to eleven_v3 (most expressive; supports [warmly]-style audio tags). If the account/MCP doesn't support v3, fall back to eleven_multilingual_v2 (no audio tags — rely on stability/style for expressiveness). State which one you're using.

Pause: "Voice: [name + id]. Model: [v3 / multilingual_v2]. Single narrator. Reply go, or tell me what to change."

Stage 3: Narration script (+ optional emotion tags)

For each scene, prepare the narration text:

  • original — the scene's text, unchanged
  • tagged — only if using eleven_v3: the same text with a few audio tags inserted to match the scene's mood. Use sparingly (1–2 tags per scene). Useful tags: [warmly], [softly], [gently], [cheerfully], [curiously], [whispering], [excited], [sadly], [reassuringly]. Place a tag BEFORE the text it affects; it persists until the next tag. Use ... for natural pauses. Don't over-tag — it reads choppy.
    • If using eleven_multilingual_v2, skip tags entirely (it ignores them / reads them aloud). Expressiveness comes from stability/style and the voice itself.
  • rationale — one line on why those tags fit (only when tagging)

Title clip (recommended). Produce a short separate title.mp3 from the bare story title with one warm tag ([warmly] The Little Cloud.). The publisher plays it on the dedicated cover page (slide 0) before auto-advancing into scene 1, so the cover isn't silent — generating it is worth the one extra clip. Keep it minimal. Save it as {slug}_audio/title.mp3.

Write the full script to {slug}_audio/narration_script.md (one ## Scene NN section per scene with the subsections above) so the user can review it in one place.

Show full SKILL.md (537 more words)Show less
Stage 4 — HARD GATE: review the narration script

Show the script and pause: "Narration script ready — review before I generate audio. Each generation is billed (~$0.10/1k chars) and non-deterministic, so a bad script wastes credits. Reply go to generate, edit for changes, or paste a corrected version."

Do NOT proceed without explicit approval. If the user wants changes: minor wording → update and re-show; "less excited / more intimate" → re-tag with the new direction; "redo scene X" → update just that scene.

Stage 5: Generate audio (one MP3 per scene)

For each scene, call mcp__elevenlabs__text_to_speech with:

  • text = the scene's tagged text (or original if not tagging)
  • voice_id = chosen voice
  • model_id = chosen model
  • stability = 0.45, style = 0.2, use_speaker_boost = true (tune to taste)
  • output_directory = stories/{slug}/{slug}_audio/

The MCP saves the file and returns its path. Rename the saved file to {slug}_part_NN.mp3 (zero-padded scene index) — use the Bash tool (mv) so filenames match the locked convention. If a generation fails, capture the error and continue with the rest of the scenes — don't crash the batch.

Leading-punctuation gotcha (Windows). The MCP derives the saved filename from the first characters of the text, so if a scene's text begins with a quote (") or other character invalid in a filename, the save fails with [Errno 22] Invalid argument even though the audio generated (and you're billed). Workaround: send that scene's text with the leading quote stripped — ElevenLabs doesn't voice quotation marks, so the spoken audio is identical. (Quotes inside the text are fine; only the first character matters for the filename.)

If you produced a title clip, generate it the same way and rename it to title.mp3.

After all scenes:

  • Write {slug}_audio/manifest.json listing each clip: index, filename, original, tagged (if any), and the scene's mood/scene metadata. Include a title_audio entry if a title clip was made.
  • Report: number of clips generated, approximate cost (total chars × $0.10/1k), any failures with the scene index for retry, and the audio folder path.
json
{
  "story_slug": "the-little-cloud",
  "voice_id": "...",
  "model_id": "eleven_v3",
  "title_audio": { "filename": "title.mp3", "original": "The Little Cloud", "tagged": "[warmly] The Little Cloud." },
  "parts": [
    { "index": 1, "filename": "the-little-cloud_part_01.mp3", "original": "High in the sky...", "tagged": "[warmly] High in the sky...", "mood": "gentle, bright" }
  ]
}
After Stage 5

The skill is done — one MP3 per scene plus a manifest, ready for the publisher. To redo specific scenes: "regenerate scenes X, Y" — generation is idempotent on filenames, so reruns overwrite cleanly. Or edit the script first and regenerate just those scenes.

Cost expectations

ElevenLabs v3 is ~$0.10 per 1,000 characters. A short beginner story (~1,500–2,500 chars across all scenes) costs ~$0.15–$0.25 once. Audio tags add ~5–10% character overhead. Budget for 1–2 regenerations of some scenes since v3 is non-deterministic — the Stage 4 gate exists to get the script right before spending on audio.

What this skill does NOT do

  • Does not edit the story text — tags are added; the prose itself doesn't change.
  • Does not generate music or sound effects — only narration.
  • Does not do voice cloning — bring a voice_id from ElevenLabs' separate workflow.
  • Does not QA the generated audio — listening and approval is the user's job.

Compatibility notes

  • Eleven v3 supports audio tags ([warmly]); eleven_multilingual_v2 does not — match your tagging to the model.
  • Eleven v3 does not reliably support SSML <break>; use ... for pauses.
  • v3 is non-deterministic — the same input can produce different audio across runs. The Stage 4 gate is the cost-control valve.
  • Filenames MUST end up as {slug}_part_NN.mp3 (and optionally title.mp3) so the publisher pairs audio with images by scene index.

© hassancs91, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/story-narrator of hassancs91/claude-image-generation.

Open the folder on GitHubat commit f533831

Compare with similar skills

Story Narrator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Story Narrator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Story Narrator this skillhassancs91/claude-image-generation102—~2.5kAutomated safety check: PassMIT
Remotion ProductionDojoCodingLabs/remotion-superpowers132—~1.1kAutomated safety check: PassMIT
Text To Speechcalesthio/OpenMontage66k—~2.4kAutomated safety check: PassAGPL-3.0
Scenario Elevenlabsscenario-labs/skills946—~2kAutomated safety check: PassMIT
Elevenlabs Agentsjezweb/claude-skills1.1k—~3.3kAutomated safety check: PassMIT
Making Demo Videosnukeop/nuclear19k—~892Automated safety check: PassAGPL-3.0

Similar skills

  • Remotion Production

    DojoCodingLabs/remotion-superpowers

    Full video production workflow for Remotion projects. An agent skill from DojoCodingLabs/remotion-superpowers.

    132 GitHub stars~1.1k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Text To Speech

    calesthio/OpenMontage

    Generate speech audio from text using HeyGen's Starfish TTS model.

    66k GitHub stars~2.4k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Scenario Elevenlabs

    scenario-labs/skills

    A skill your agent uses when generating or transforming audio with ElevenLabs models on Scenario via MCP: text-to-speech with inline emotion tags, music from a prompt or sectioned songs, sound…

    946 GitHub stars~2k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Elevenlabs Agents

    jezweb/claude-skills

    Build conversational AI voice agents on the ElevenLabs platform.

    1.1k GitHub stars~3.3k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Making Demo Videos

    nukeop/nuclear

    A skill your agent uses when making a demo, tutorial, or feature video of Nuclear.

    19k GitHub stars~892 tokensUpdated 4 days ago
    Media & CreativeAuto-check passed
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed

More from hassancs91/claude-image-generation

  • 3D Image Renderer

    hassancs91/claude-image-generation

    Generate PNG images by building a real Three.js 3D scene and capturing one frame headlessly — no image model involved.

    102 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Story HTML Publisher

    hassancs91/claude-image-generation

    Final step of the AI Storybook pipeline. An agent skill from hassancs91/claude-image-generation.

    102 GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Cf Image

    hassancs91/claude-image-generation

    Generate an image from a text description using Cloudflare Workers AI (the flux-1-schnell model).

    102 GitHub stars~867 tokensUpdated 1 mo ago
    Auto-check: notes
  • Prompt To Design

    hassancs91/claude-image-generation

    Generate a polished PNG graphic from a text prompt and an aspect ratio.

    102 GitHub stars~3.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Scene Splitter

    hassancs91/claude-image-generation

    Splits a plain English story into a numbered list of SCENES — each scene being one moment that gets exactly one illustration AND one narration clip downstream.

    102 GitHub stars~2.3k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Story Narrator

What does Story Narrator do?

Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech). Story Narrator is an agent skill from hassancs91/claude-image-generation. Generates one expressive narration MP3 per scene for an English storybook, using the ElevenLabs MCP (texttospeech).

When should I use Story Narrator?

Story Narrator fits situations like: the user wants to narrate a story; generate per-scene audio; voice a storybook; create read-along narration.

How do I install Story Narrator in Claude Code?

Run `npx skills add hassancs91/claude-image-generation --skill story-narrator -a claude-code`. Or copy the skill folder (.claude/skills/story-narrator in hassancs91/claude-image-generation) into .claude/skills/story-narrator in your project. Claude Code loads it when a task matches its description.

How do I install Story Narrator in Codex?

Run `npx skills add hassancs91/claude-image-generation --skill story-narrator -a codex`. Or copy the skill folder (.claude/skills/story-narrator in hassancs91/claude-image-generation) into .agents/skills/story-narrator in your project. Codex loads it when a task matches its description.

Can I use Story Narrator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hassancs91/claude-image-generation --skill story-narrator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/story-narrator, .gemini/skills/story-narrator, .github/skills/story-narrator and .opencode/skills/story-narrator in your project.

What does Story Narrator need to run?

Going by SKILL.md and its folder, Story Narrator needs the command-line tools its instructions call (uvx) and credentials named ELEVENLABS_API_KEY. Our summary lists: A credential in ELEVENLABS_API_KEY.

Does Story Narrator access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Story Narrator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Story Narrator use?

Story Narrator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Story Narrator use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Story Narrator?

Skills that share tags, products or a category with Story Narrator: Remotion Production (DojoCodingLabs/remotion-superpowers, 132 stars), Text To Speech (calesthio/OpenMontage, 66k stars), Scenario Elevenlabs (scenario-labs/skills, 946 stars) and Elevenlabs Agents (jezweb/claude-skills, 1.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Story Narrator?

hassancs91 (a GitHub user) maintains it in hassancs91/claude-image-generation, which has 102 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on August 18, 2026.

Source: hassancs91/claude-image-generation on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.