Agent skill

Speech

by davila7 in davila7/claude-code-templates

A skill your agent uses when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation.

Apache-2.0Auto-check passedMedia & Creative

Install Speech

skills CLI
$ npx skills add davila7/claude-code-templates --skill speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/media/speech .claude/skills/speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speech
GitHub stars
32k
Token cost
~2k tokens
SKILL.md length
854 words
Files
20 (incl. scripts, references, assets)
Skills in repo
477
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation.

  • Works in 8 steps: Decide intent: single vs batch (see… → Collect inputs up front: exact text… → If batch: write a temporary JSONL under… → …
  • The user asks for text-to-speech narration
  • SKILL.md covers When to use, Decision tree (single vs batch), Workflow and Temp and output conventions, plus 9 more sections
  • Runs Python scripts from its folder; calls uv and python3; needs OPENAI_API_KEY and ATLASCLOUD_API_KEY

What it does

Speech is an agent skill from davila7/claude-code-templates. Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation. OpenAI remains the default; Atlas Cloud is an explicit optional backend for asynchronous multilingual speech. Custom voice creation is out of scope.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 23 other files, including scripts, reference files and assets (for example `agents/openai.yaml`, `references/accessibility.md` and `references/atlas-cloud.md`).

It sits in Media & Creative, covering Text to speech and voice. It works with OpenAI. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is Apache-2.0.

When your agent uses it

  • The user asks for text-to-speech narration
  • Accessibility reads
  • Batch speech generation

Example prompts

  • “/speech”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • A credential in ATLASCLOUD_API_KEY

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Decide intent: single vs batch (see decision tree above).
  2. Collect inputs up front: exact text (verbatim), desired voice, delivery style, format, and any constraints.
  3. If batch: write a temporary JSONL under tmp/ (one job per line), run once, then delete the JSONL.
  4. Augment instructions into a short labeled spec without rewriting the input text.
  5. Run scripts/text_to_speech.py for the default OpenAI path, or scripts/atlas_text_to_speech.py only when Atlas Cloud was selected (see…
  6. For important clips, validate: intelligibility, pacing, pronunciation, and adherence to constraints.
  7. Iterate with a single targeted change (voice, speed, or instructions), then re-check.
  8. Save/return final outputs and note the final text + instructions + flags used.

What it can do on your machine

Read from SKILL.md and the folder at commit 4c82aba. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ATLASCLOUD_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Speech loads about 2k tokens when it runs, and up to ~6.1k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 854 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 4c82aba, republished under its Apache-2.0 licence (© davila7). 854 words, ~2,004 tokens.

Download SKILL.mdSave it as .claude/skills/speech/SKILL.md (or your agent's skills folder). This skill also uses 19 other files; get the full folder from GitHub.
name
speech
description
Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation. OpenAI remains the default; Atlas Cloud is an explicit optional backend for asynchronous multilingual speech. Custom voice creation is out of scope.
author
openai

Speech Generation Skill

Generate spoken audio for the current project (narration, product demo voiceover, IVR prompts, accessibility reads). OpenAI remains the default with gpt-4o-mini-tts-2025-12-15; Atlas Cloud is available only when the user explicitly selects it. Prefer the bundled CLIs for deterministic, reproducible runs.

When to use

  • Generate a single spoken clip from text
  • Generate a batch of prompts (many lines, many files)

Decision tree (single vs batch)

  • If the user provides multiple lines/prompts or wants many outputs -> batch
  • Else -> single

Workflow

  1. Decide intent: single vs batch (see decision tree above).
  2. Collect inputs up front: exact text (verbatim), desired voice, delivery style, format, and any constraints.
  3. If batch: write a temporary JSONL under tmp/ (one job per line), run once, then delete the JSONL.
  4. Augment instructions into a short labeled spec without rewriting the input text.
  5. Run scripts/text_to_speech.py for the default OpenAI path, or scripts/atlas_text_to_speech.py only when Atlas Cloud was selected (see references/cli.md).
  6. For important clips, validate: intelligibility, pacing, pronunciation, and adherence to constraints.
  7. Iterate with a single targeted change (voice, speed, or instructions), then re-check.
  8. Save/return final outputs and note the final text + instructions + flags used.

Temp and output conventions

  • Use tmp/speech/ for intermediate files (for example JSONL batches); delete when done.
  • Write final artifacts under output/speech/ when working in this repo.
  • Use --out or --out-dir to control output paths; keep filenames stable and descriptive.

Dependencies (install if missing)

Prefer uv for dependency management.

OpenAI backend package:

uv pip install openai

If uv is unavailable:

python3 -m pip install openai

The Atlas Cloud backend uses only the Python standard library.

Environment

  • Default OpenAI calls require OPENAI_API_KEY.
  • Optional Atlas Cloud calls require ATLASCLOUD_API_KEY.

If the selected provider key is missing, give the user these steps:

  1. Create an API key in that provider's console.
  2. Set OPENAI_API_KEY or ATLASCLOUD_API_KEY as an environment variable in their system.
  3. Offer to guide them through setting the environment variable for their OS/shell if needed.
  • Keep provider credentials out of chat. Direct the user to set the selected key locally and confirm when ready.

If installation isn't possible in this environment, tell the user which dependency is missing and how to install it locally.

Defaults & rules

  • Keep OpenAI as the default provider. Do not switch to Atlas Cloud unless the user requests it.
  • Use gpt-4o-mini-tts-2025-12-15 unless the user requests another model.
  • Default voice: cedar. If the user wants a brighter tone, prefer marin.
  • Built-in voices only. Custom voices are out of scope for this skill.
  • instructions are supported for GPT-4o mini TTS models, but not for tts-1 or tts-1-hd.
  • Input length must be <= 4096 characters per request. Split longer text into chunks.
  • Enforce 50 requests/minute. The CLI caps --rpm at 50.
  • Require OPENAI_API_KEY before any live API call.
  • For Atlas Cloud, default to xai/tts-v1, voice eve, language auto, and require ATLASCLOUD_API_KEY.
  • Atlas generation submits exactly one POST. Only prediction GET requests may retry, and polling must stay finite.
  • Download Atlas outputs without an Authorization header and reject non-HTTPS or private-network targets.
  • Provide a clear disclosure to end users that the voice is AI-generated.
  • Use the OpenAI Python SDK (openai package) for default OpenAI calls; the dedicated Atlas CLI uses its asynchronous HTTP contract.
  • Prefer the matching bundled CLI over writing new one-off scripts.
Show full SKILL.md (317 more words)Show less

Instruction augmentation

Reformat user direction into a short, labeled spec. Only make implicit details explicit; do not invent new requirements.

Quick clarification (augmentation vs invention):

  • If the user says "narration for a demo", you may add implied delivery constraints (clear, steady pacing, friendly tone).
  • Do not introduce a new persona, accent, or emotional style the user did not request.

Template (include only relevant lines):

Voice Affect: <overall character and texture of the voice>
Tone: <attitude, formality, warmth>
Pacing: <slow, steady, brisk>
Emotion: <key emotions to convey>
Pronunciation: <words to enunciate or emphasize>
Pauses: <where to add intentional pauses>
Emphasis: <key words or phrases to stress>
Delivery: <cadence or rhythm notes>

Augmentation rules:

  • Keep it short; add only details the user already implied or provided elsewhere.
  • Do not rewrite the input text.
  • If any critical detail is missing and blocks success, ask a question; otherwise proceed.

Examples

Single example (narration)
Input text: "Welcome to the demo. Today we'll show how it works."
Instructions:
Voice Affect: Warm and composed.
Tone: Friendly and confident.
Pacing: Steady and moderate.
Emphasis: Stress "demo" and "show".
Batch example (IVR prompts)
{"input":"Thank you for calling. Please hold.","voice":"cedar","response_format":"mp3","out":"hold.mp3"}
{"input":"For sales, press 1. For support, press 2.","voice":"marin","instructions":"Tone: Clear and neutral. Pacing: Slow.","response_format":"wav"}

Instructioning best practices (short list)

  • Structure directions as: affect -> tone -> pacing -> emotion -> pronunciation/pauses -> emphasis.
  • Keep 4 to 8 short lines; avoid conflicting guidance.
  • For names/acronyms, add pronunciation hints (e.g., "enunciate A-I") or supply a phonetic spelling in the text.
  • For edits/iterations, repeat invariants (e.g., "keep pacing steady") to reduce drift.
  • Iterate with single-change follow-ups.

More principles: references/prompting.md. Copy/paste specs: references/sample-prompts.md.

Guidance by use case

Use these modules when the request is for a specific delivery style. They provide targeted defaults and templates.

  • Narration / explainer: references/narration.md
  • Product demo / voiceover: references/voiceover.md
  • IVR / phone prompts: references/ivr.md
  • Accessibility reads: references/accessibility.md

CLI + environment notes

  • CLI commands + examples: references/cli.md
  • API parameter quick reference: references/audio-api.md
  • Atlas Cloud asynchronous workflow and safeguards: references/atlas-cloud.md
  • Instruction patterns + examples: references/voice-directions.md
  • If network approvals / sandbox settings are getting in the way: references/codex-network.md

Reference map

  • references/cli.md: how to run speech generation/batches via scripts/text_to_speech.py (commands, flags, recipes).
  • references/audio-api.md: API parameters, limits, voice list.
  • references/atlas-cloud.md: optional Atlas Cloud model, CLI, polling, and download contract.
  • references/voice-directions.md: instruction patterns and examples.
  • references/prompting.md: instruction best practices (structure, constraints, iteration patterns).
  • references/sample-prompts.md: copy/paste instruction recipes (examples only; no extra theory).
  • references/narration.md: templates + defaults for narration and explainers.
  • references/voiceover.md: templates + defaults for product demo voiceovers.
  • references/ivr.md: templates + defaults for IVR/phone prompts.
  • references/accessibility.md: templates + defaults for accessibility reads.
  • references/codex-network.md: environment/sandbox/network-approval troubleshooting.

© davila7, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 19 other files (scripts, references, assets) in cli-tool/components/skills/media/speech of davila7/claude-code-templates.

  • SKILL.md
  • LICENSE.txt
  • agents/openai.yaml
  • assets/speech-small.svg
  • assets/speech.png
  • references/accessibility.md
  • references/atlas-cloud.md
  • references/audio-api.md
  • references/cli.md
  • references/codex-network.md
  • references/ivr.md
  • references/narration.md
  • references/prompting.md
  • references/sample-prompts.md
  • references/voice-directions.md
  • references/voiceover.md
  • scripts/atlas_text_to_speech.py
  • … and 3 more

Open the folder on GitHubat commit 4c82aba

Compare with similar skills

Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speech this skilldavila7/claude-code-templates32k—~2kAutomated safety check: PassApache-2.0
SpeechJetBrains/skills3634 repos~1.9kAutomated safety check: PassApache-2.0
Podcastteam-attention/plugins-for-claude-natives825—~1.5kAutomated safety check: PassMIT
Lessongug007/lpm152—~1.2kAutomated safety check: PassMIT
Voxclawmalpern/VoxClaw208—~1.9kAutomated safety check: PassNone
Speechnexu-io/open-design100k—~291Automated safety check: PassApache-2.0

Similar skills

  • Speech

    JetBrains/skills

    Official

    A skill your agent uses when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation via the OpenAI Audio API; run the bundled CLI…

    363 GitHub starsUsed in 4 repos~1.9k tokens
    Media & CreativeAuto-check passed
  • Podcast

    team-attention/plugins-for-claude-natives

    Generate Korean podcast episodes from any source (URLs, tweets, articles, PDFs) — analyzes content, writes a script, generates audio via OpenAI TTS, converts to MP4, and auto-uploads to YouTube.

    825 GitHub stars~1.5k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Lesson

    gug007/lpm

    Make a narrated lesson video about lpm from the real desktop app, recorded on a pristine data directory.

    152 GitHub stars~1.2k tokensUpdated today
    Media & CreativeAuto-check passed
  • Voxclaw

    malpern/VoxClaw

    Give your agent a voice. An agent skill from malpern/VoxClaw.

    208 GitHub stars~1.9k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Speech

    nexu-io/open-design

    Generate spoken audio from text using OpenAI's API with built-in voices.

    100k GitHub stars~291 tokensUpdated today
    Media & CreativeAuto-check passed
  • Openai Tts

    benchflow-ai/skillsbench

    OpenAI Text-to-Speech API for high-quality speech synthesis.

    1.8k GitHub stars~946 tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed

More from davila7/claude-code-templates

All 477 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 12 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 10 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 2 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Works with

Questions about Speech

What does Speech do?

A skill your agent uses when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation. Speech is an agent skill from davila7/claude-code-templates. Use when the user asks for text-to-speech narration or voiceover, accessibility reads, audio prompts, or batch speech generation.

When should I use Speech?

Speech fits situations like: the user asks for text-to-speech narration; accessibility reads; batch speech generation.

How do I install Speech in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill speech -a claude-code`. Or copy the skill folder (cli-tool/components/skills/media/speech in davila7/claude-code-templates) into .claude/skills/speech in your project. Claude Code loads it when a task matches its description.

How do I install Speech in Codex?

Run `npx skills add davila7/claude-code-templates --skill speech -a codex`. Or copy the skill folder (cli-tool/components/skills/media/speech in davila7/claude-code-templates) into .agents/skills/speech in your project. Codex loads it when a task matches its description.

Can I use Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech, .gemini/skills/speech, .github/skills/speech and .opencode/skills/speech in your project.

What does Speech need to run?

Going by SKILL.md and its folder, Speech needs Python for the scripts in its folder, the command-line tools its instructions call (uv and python3) and credentials named OPENAI_API_KEY and ATLASCLOUD_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in ATLASCLOUD_API_KEY.

Does Speech access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Speech safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Speech use?

Speech is published under the Apache-2.0 licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Speech use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.1k tokens, read only when the agent opens those files.

What are the alternatives to Speech?

Skills that share tags, products or a category with Speech: Speech (JetBrains/skills, 363 stars), Podcast (team-attention/plugins-for-claude-natives, 825 stars), Lesson (gug007/lpm, 152 stars) and Voxclaw (malpern/VoxClaw, 208 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speech?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,432 GitHub stars. The repository holds 477 skills in this directory. The repository was last updated on October 7, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.