Agent skill

Fish Audio Tts

by calesthio in calesthio/OpenMontage

Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via referenceid.

AGPL-3.0Auto-check: notesMedia & Creative

Install Fish Audio Tts

skills CLI
$ npx skills add calesthio/OpenMontage --skill fish-audio-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage fish-audio-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/fish-audio-tts .claude/skills/fish-audio-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fish-audio-tts
GitHub stars
66k
Token cost
~1.4k tokens
SKILL.md length
574 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via referenceid.

  • Works in 4 steps: Generate a 10-15 second sample with the… → Ask the user to approve voice… → Generate the full narration only after… → …
  • The user prefers fish.audio/Fish Audio TTS
  • SKILL.md covers Current API, Backend models, Inline emotion tags (S2 models… and Voice selection (reference_id), plus 5 more sections
  • Reaches api.fish.audio; needs FISH_AUDIO_API_KEY

What it does

Fish Audio Tts is an agent skill from calesthio/OpenMontage. Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via referenceid. Use when the user prefers fish.audio/Fish Audio TTS, wants a specific playground voice model, or needs high-emotion voice-clone narration.

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is AGPL-3.0.

When your agent uses it

  • The user prefers fish.audio/Fish Audio TTS
  • Wants a specific playground voice model
  • Needs high-emotion voice-clone narration

Example prompts

  • “/fish-audio-tts”

Requirements

  • Python 3
  • A credential in FISH_AUDIO_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Generate a 10-15 second sample with the chosen model + reference_id before a full paid narration.
  2. Ask the user to approve voice naturalness, emotion, and pace.
  3. Generate the full narration only after approval.
  4. For batch/localization variants where cost matters, prototype on s2.1-pro-free (promo-window $0; non-commercial drafts only) and upgrade…

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.fish.audio

    Also links to:

    • fish.audio
    • docs.fish.audio

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FISH_AUDIO_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fish Audio Tts loads about 1.4k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 574 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:8
    Requires `FISH_AUDIO_API_KEY` in `.env` (create one at https://fish.audio/go-api/api-keys/).

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its AGPL-3.0 licence (© calesthio). 574 words, ~1,361 tokens.

Download SKILL.mdSave it as .claude/skills/fish-audio-tts/SKILL.md (or your agent's skills folder).
name
fish-audio-tts
description
Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via reference_id. Use when the user prefers fish.audio/Fish Audio TTS, wants a specific playground voice model, or needs high-emotion voice-clone narration.

fish.audio TTS

Requires FISH_AUDIO_API_KEY in .env (create one at https://fish.audio/go-api/api-keys/). Create voice models in the fish.audio playground and pass their id as reference_id to reuse a cloned voice.

Current API

Single synchronous call returning raw audio bytes:

text
POST https://api.fish.audio/v1/tts
Authorization: Bearer ${FISH_AUDIO_API_KEY}
Content-Type: application/json
model: <backend model>     # HTTP header selects the backend, e.g. s1

The backend model is chosen with the model HTTP header, not a body field. In OpenMontage this maps to the tool's model input.

Backend models

model is required — there is no default. Pass one of:

  • s2.1-pro — latest generation. Best quality: inline emotion tags, 80+ languages, multi-speaker. Hero narration.
  • s2.1-pro-free — promotional free access to s2.1-pro. Drafts, samples, and validation runs at $0 during the promo window only. Per the fish.audio announcement: free through August 31, 2026, subject to Fair Use, no SLA/latency guarantee, requests may be retained, and commercial use is restricted. Never route production or client narration through it.
  • s2-pro — first S2 generation. Stable high quality with emotion-tag support.
  • s1 — previous flagship. Kept for compatibility with existing integrations.

Billing is per UTF-8 byte of input text (not per character). CJK text and emoji cost 3-4x an ASCII character of the same visible length. Current list pricing: s1 / s2-pro / s2.1-pro = $15 per 1M bytes, s2.1-pro-free = $0 during the promo window only (the tool's estimate_cost() switches to the paid s2.1-pro rate after August 31, 2026). Verify current pricing at https://docs.fish.audio/developer-guide/models-pricing/pricing-and-rate-limits before large batches.

Inline emotion tags (S2 models only)

s2-pro / s2.1-pro / s2.1-pro-free interpret inline emotion tags embedded in the text:

  • Tags like [laugh], [whispers] change the delivery mid-sentence.
  • Example: "That's hilarious [laugh] but let me explain seriously."
  • s1 does not interpret emotion tags — they may be read out as plain text, so strip them when targeting s1.

Voice selection (reference_id)

  • Build or pick a voice in the fish.audio playground, then copy its model id.
  • Pass it as reference_id. The selector's generic voice_id is accepted as an alias when reference_id is absent.
  • Without a reference_id, fish.audio uses its default voice for the chosen model.

Inline on-the-fly cloning (uploading reference audio + text per request) is not supported by this tool — create a voice model in the playground first.

Show full SKILL.md (233 more words)Show less

OpenMontage Usage

Generate with the TTS selector:

python
from tools.audio.tts_selector import TTSSelector

result = TTSSelector().execute({
    "preferred_provider": "fish_audio",
    "text": "Here's why compound interest quietly beats every get-rich-quick scheme.",
    "model": "s1",
    "reference_id": "<playground voice model id>",
    "output_path": "projects/my-video/assets/audio/narration.mp3",
})

Or call the provider directly:

python
from tools.audio.fish_audio_tts import FishAudioTTS

result = FishAudioTTS().execute({
    "text": "Short sample line for approval.",
    "model": "s1",
    "reference_id": "<playground voice model id>",
    "output_path": "projects/my-video/assets/audio/fish_sample.mp3",
})

The provider writes the audio to output_path and returns data.output plus the resolved model and reference_id.

Quality & latency tuning

  • latency: normal (default, best quality), balanced (a little faster), or low (fastest, slight quality cost).
  • normalize: default true; keep it on so numbers, dates, and currency read naturally.
  • prosody: optional { "speed": 1.0, "volume": 0 } to nudge pace/loudness.
  • mp3_bitrate: 128 is a good default; raise to 192 for music-bed-heavy mixes.
  • temperature: default 0.7. Raise toward 0.9 for more expressive reads (recommended when leaning on emotion tags); lower for a steadier, more predictable delivery.
  • top_p / repetition_penalty: usually leave at the defaults (0.7 / 1.2).
  1. Generate a 10-15 second sample with the chosen model + reference_id before a full paid narration.
  2. Ask the user to approve voice naturalness, emotion, and pace.
  3. Generate the full narration only after approval.
  4. For batch/localization variants where cost matters, prototype on s2.1-pro-free (promo-window $0; non-commercial drafts only) and upgrade the final to s2.1-pro.

Troubleshooting

  • 401 Unauthorized: wrong or missing FISH_AUDIO_API_KEY.
  • 402 / payment errors: account credit exhausted.
  • 404 / bad voice: the reference_id is wrong or not owned by this account.
  • Empty/short audio: check that text is non-empty and normalize is not stripping the whole input.

Safety

Never print or write the API key to logs, metadata, patches, or project artifacts. .env.example should contain only empty variable names.

© calesthio, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/fish-audio-tts of calesthio/OpenMontage.

Open the folder on GitHubat commit 9327439

Compare with similar skills

Fish Audio Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fish Audio Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fish Audio Tts this skillcalesthio/OpenMontage66k—~1.4kAutomated safety check: NotesAGPL-3.0
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo130k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    130k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from calesthio/OpenMontage

All 41 skills in this repo
  • Video Understand

    calesthio/OpenMontage

    Understand video content locally using ffmpeg frame extraction and Whisper transcription.

    66k GitHub stars~841 tokensUpdated 7 days ago
    Auto-check passed
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    66k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    66k GitHub stars~2k tokensUpdated 7 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    66k GitHub stars~855 tokensUpdated 7 days ago
    Auto-check: notes

Questions about Fish Audio Tts

What does Fish Audio Tts do?

Generate expressive, multilingual narration with fish.audio (S1 / S2-generation models) and reuse cloned voices via referenceid. Fish Audio Tts is an agent skill from calesthio/OpenMontage.audio (S1 / S2-generation models) and reuse cloned voices via referenceid.

When should I use Fish Audio Tts?

Fish Audio Tts fits situations like: the user prefers fish.audio/Fish Audio TTS; wants a specific playground voice model; needs high-emotion voice-clone narration.

How do I install Fish Audio Tts in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill fish-audio-tts -a claude-code`. Or copy the skill folder (.agents/skills/fish-audio-tts in calesthio/OpenMontage) into .claude/skills/fish-audio-tts in your project. Claude Code loads it when a task matches its description.

How do I install Fish Audio Tts in Codex?

Run `npx skills add calesthio/OpenMontage --skill fish-audio-tts -a codex`. Or copy the skill folder (.agents/skills/fish-audio-tts in calesthio/OpenMontage) into .agents/skills/fish-audio-tts in your project. Codex loads it when a task matches its description.

Can I use Fish Audio Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill fish-audio-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fish-audio-tts, .gemini/skills/fish-audio-tts, .github/skills/fish-audio-tts and .opencode/skills/fish-audio-tts in your project.

What does Fish Audio Tts need to run?

Going by SKILL.md and its folder, Fish Audio Tts needs credentials named FISH_AUDIO_API_KEY. Our summary lists: Python 3; A credential in FISH_AUDIO_API_KEY.

Does Fish Audio Tts access the network?

SKILL.md names 3 domains. In commands or code: api.fish.audio; the agent is likely to contact it when it follows the instructions. As links in the text: fish.audio and docs.fish.audio. This is read from the text; nothing was executed.

Is Fish Audio Tts safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Fish Audio Tts use?

Fish Audio Tts is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fish Audio Tts use?

About 1.4k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fish Audio Tts?

Skills that share tags, products or a category with Fish Audio Tts: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 130k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fish Audio Tts?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.