Agent skill

Fish Audio API

by waker240 in waker240/FullVideoProductionSkill

Generate narration with the public Fish Audio REST API using environment credentials and a user-selected voice.

Apache-2.0Auto-check: notesMedia & Creative

Install Fish Audio API

skills CLI
$ npx skills add waker240/FullVideoProductionSkill --skill fish-audio-api -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install waker240/FullVideoProductionSkill fish-audio-api --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/waker240/FullVideoProductionSkill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/fish-audio-api .claude/skills/fish-audio-api && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
fish-audio-api
GitHub stars
197
Token cost
~1.1k tokens
SKILL.md length
468 words
Files
5 (incl. scripts)
Skills in repo
13
Repo updated
First seen
Licence
Apache-2.0

At a glance

Generate narration with the public Fish Audio REST API using environment credentials and a user-selected voice.

  • Tasks that involve Text to speech and voice
  • SKILL.md covers Configure locally, Generate and revise, Optional public media APIs and Maintainer verification
  • Runs JavaScript scripts from its folder; calls node; reaches api.fish.audio; needs FISH_API_KEY
  • Tasks that involve Image generation

What it does

Fish Audio API is an agent skill from waker240/FullVideoProductionSkill. Generate narration with the public Fish Audio REST API using environment credentials and a user-selected voice. Provides shared public API helpers for Fish narration, OpenAI word timestamps, and optional GPT Image generation; no browser cookies or private proxy.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts.

It sits in Media & Creative, covering Text to speech and voice, Image generation and REST APIs. It works with OpenAI. The repository describes itself as: 20亿Claude Opus 4.6-4.8的Token. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Text to speech and voice
  • Tasks that involve Image generation
  • Tasks that involve REST APIs

Example prompts

  • “/fish-audio-api”

Requirements

  • Node.js
  • A credential in FISH_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 0223baa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (JavaScript), which the agent can run.

    Shell commands in SKILL.md call:

    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.fish.audio

    Also links to:

    • docs.fish.audio

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • FISH_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Fish Audio API loads about 1.1k tokens when it runs. Until then it costs about 69 tokens; SKILL.md has 468 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~69
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:12
    env.example) to the **video project's** `.env`, ensure `.env` is ignored by Git, then fill it locally. Never request tha

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from waker240/FullVideoProductionSkill at commit 0223baa, republished under its Apache-2.0 licence (© waker240). 468 words, ~1,075 tokens.

Download SKILL.mdSave it as .claude/skills/fish-audio-api/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
fish-audio-api
description
Generate narration with the public Fish Audio REST API using environment credentials and a user-selected voice. Provides shared public API helpers for Fish narration, OpenAI word timestamps, and optional GPT Image generation; no browser cookies or private proxy.

Fish Audio public API

Use this skill when a video needs Fish narration. The release uses the official POST https://api.fish.audio/v1/tts endpoint. It does not ship a web-session proxy, cookies, an account, or a voice. Node.js 22.20 or later is required. Install FFmpeg/ffprobe for the video audio pipeline.

Configure locally

Copy env.example to the video project's .env, ensure .env is ignored by Git, then fill it locally. Never request that users paste credentials into chat. Existing process environment values take precedence over .env. Scripts never print secret values or upstream error bodies.

VariableMeaning
FISH_API_KEYAPI key created in the user's Fish account; required for generation.
FISH_REFERENCE_IDA voice ID the user has permission to use; required unless a project explicitly supplies one.
FISH_MODELDefaults to s2.1-pro-free; supported choices also include s1, s2-pro, s2.1-pro. Account access and quotas are provider-controlled.

No silent model fallback is allowed. An unknown model fails locally, because the provider may otherwise choose a paid default. Availability and pricing can change; check the account before generation. Setting a key does not itself authorize a charge: run synthesis only for the user's requested scope.

Generate and revise

For a scaffolded HyperFrames project, edit scripts/narration.json. Its voice.referenceId and voice.model may be null to use the environment. All other voice settings are explicit in the template. Preview one short paragraph before producing a whole narration.

sh
node scripts/tts-fish.mjs --section s0
node scripts/tts-fish.mjs
node scripts/tts-fish.mjs --section s2 --para 1 --force

--para requires every unselected paragraph in that section to exist. To start a new project, first generate one short section with --section s0; use --para for a later repair. Matching paragraph text and voice profiles reuse frozen audio. --force makes another API request for the selected speech.

The common helper public-media-api.cjs supports the shared audio engine as well. Its Fish configuration accepts fish.transport: "official-api", fish.reference_id, fish.model, and fish.request. Sampling fields are top-level request fields, such as temperature and top_p; a website-style backend/sampler profile is not this API contract. Never put credentials in JSON project profiles.

Show full SKILL.md (148 more words)Show less

Every generation POST is sent once. A timeout can occur after the provider started work; inspect provider usage before deliberately retrying. Redirects are refused. Official endpoints are pinned. FISH_API_URL with HYPERFRAMES_TEST_ALLOW_FISH_LOOPBACK=1 is only for local mock tests.

Optional public media APIs

The same helper supports OpenAI whisper-1 transcription with word timestamps, and GPT Image PNG generation. These are separate account/billing capabilities, not features bundled by an agent subscription. See requirements and transcription. Images require an explicit OPENAI_IMAGE_MODEL or --model; choose a GPT Image model available to the user's account. External image/video tools, including Codex image tools and video-generation plugins, are optional and are not installed or authenticated by this skill package.

Maintainer verification

sh
node --test skills/fish-audio-api/tests/public-media-api.test.cjs
node --test skills/hyperframes/scripts/tests/fish-tts.test.mjs

These tests use fake keys and loopback HTTP, not paid services. The release verification does not establish that a user's account has model access.

Official references checked on 2026-09-26: Fish TTS API, Fish developer guide.

© waker240, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in skills/fish-audio-api of waker240/FullVideoProductionSkill.

  • SKILL.md
  • .gitignore
  • env.example
  • scripts/public-media-api.cjs
  • tests/public-media-api.test.cjs

Open the folder on GitHubat commit 0223baa

Compare with similar skills

Fish Audio API next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Fish Audio API compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Fish Audio API this skillwaker240/FullVideoProductionSkill197—~1.1kAutomated safety check: NotesApache-2.0
Local AI Useamd/skills398—~5kAutomated safety check: NotesMIT
Agentselevenlabs/skills481—~6.5kAutomated safety check: PassMIT
Local AI App Integrationamd/skills398—~6kAutomated safety check: PassMIT
Codex Imagendarkamenosa/codex-imagen135—~2.6kAutomated safety check: PassMIT
Chatgpt CLIItamarZand88/CLI-Anything-WEB231—~781Automated safety check: PassMIT

Similar skills

  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    398 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Agents

    elevenlabs/skills

    Build voice AI agents with ElevenLabs. An agent skill from elevenlabs/skills.

    481 GitHub stars~6.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    398 GitHub stars~6k tokensUpdated today
    Media & CreativeAuto-check passed
  • Codex Imagen

    darkamenosa/codex-imagen

    Generate or edit raster images by calling the ChatGPT/Codex hosted imagegeneration flow with local Codex or OpenClaw OAuth credentials, then save decoded image files for OpenClaw and other agent…

    135 GitHub stars~2.6k tokensUpdated 14 days ago
    Media & CreativeAuto-check passed
  • Chatgpt CLI

    ItamarZand88/CLI-Anything-WEB

    Drives ChatGPT from the terminal via cli-web-chatgpt — ask questions, generate and download images, list and view conversations, browse models, and manage OpenAI SSO auth.

    231 GitHub stars~781 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Gpt Image Gen

    ninehills/skills

    生图 / 生成图片 / 画图 — 用 OpenAI gpt-image-2 生成图像。支持文生图、参考图生图 (img2img)、蒙版修补 (inpainting)。当用户要求用 GPT 画图、OpenAI 生图、gpt-image-2、文+图生图、参考图片生成、img2img、inpainting 时必加载此技能。Auth 自动继承 OPENAIAPIKEY / Codex OAuth…

    281 GitHub stars~1.8k tokensUpdated 3 mo ago
    Media & CreativeAuto-check: notes

More from waker240/FullVideoProductionSkill

All 13 skills in this repo
  • Hyperframes Animation

    waker240/FullVideoProductionSkill

    Build deterministic animation for HyperFrames using motion rules, transitions, scene blueprints, and runtime adapters.

    197 GitHub stars~2.4k tokensUpdated 11 days ago
    Auto-check passed
  • Hyperframes Creative

    waker240/FullVideoProductionSkill

    Develop video direction, typography, color, visual mechanisms, storyboards, and asset prompts.

    197 GitHub stars~3k tokensUpdated 11 days ago
    Auto-check passed
  • Media Use

    waker240/FullVideoProductionSkill

    Resolve reviewed local BGM, sound effects, images, icons, and brand assets into frozen project files plus a manifest.

    197 GitHub stars~2k tokensUpdated 11 days ago
    Auto-check: notes
  • Music To Video

    waker240/FullVideoProductionSkill

    Build a HyperFrames video around a supplied music track, using measured rhythm and editable visual concepts.

    197 GitHub stars~4.3k tokensUpdated 11 days ago
    Auto-check passed
  • Remotion To Hyperframes

    waker240/FullVideoProductionSkill

    Port an existing Remotion composition to HyperFrames HTML while preserving visible behavior.

    197 GitHub stars~2.4k tokensUpdated 11 days ago
    Auto-check passed
  • Hyperframes Media

    waker240/FullVideoProductionSkill

    Prepare Fish official-API narration, OpenAI word timestamps, and reviewed user-supplied local music/SFX for HyperFrames.

    197 GitHub stars~1.2k tokensUpdated 11 days ago
    Auto-check: notes

Works with

Questions about Fish Audio API

What does Fish Audio API do?

Generate narration with the public Fish Audio REST API using environment credentials and a user-selected voice. Fish Audio API is an agent skill from waker240/FullVideoProductionSkill. Generate narration with the public Fish Audio REST API using environment credentials and a user-selected voice.

When should I use Fish Audio API?

Fish Audio API fits situations like: tasks that involve Text to speech and voice; tasks that involve Image generation; tasks that involve REST APIs.

How do I install Fish Audio API in Claude Code?

Run `npx skills add waker240/FullVideoProductionSkill --skill fish-audio-api -a claude-code`. Or copy the skill folder (skills/fish-audio-api in waker240/FullVideoProductionSkill) into .claude/skills/fish-audio-api in your project. Claude Code loads it when a task matches its description.

How do I install Fish Audio API in Codex?

Run `npx skills add waker240/FullVideoProductionSkill --skill fish-audio-api -a codex`. Or copy the skill folder (skills/fish-audio-api in waker240/FullVideoProductionSkill) into .agents/skills/fish-audio-api in your project. Codex loads it when a task matches its description.

Can I use Fish Audio API in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add waker240/FullVideoProductionSkill --skill fish-audio-api -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/fish-audio-api, .gemini/skills/fish-audio-api, .github/skills/fish-audio-api and .opencode/skills/fish-audio-api in your project.

What does Fish Audio API need to run?

Going by SKILL.md and its folder, Fish Audio API needs JavaScript for the scripts in its folder, the command-line tools its instructions call (node) and credentials named FISH_API_KEY. Our summary lists: Node.js; A credential in FISH_API_KEY.

Does Fish Audio API access the network?

SKILL.md names 2 domains. In commands or code: api.fish.audio; the agent is likely to contact it when it follows the instructions. As links in the text: docs.fish.audio. This is read from the text; nothing was executed.

Is Fish Audio API safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Fish Audio API use?

Fish Audio API is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Fish Audio API use?

About 1.1k tokens (SKILL.md is roughly 4.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Fish Audio API?

Skills that share tags, products or a category with Fish Audio API: Local AI Use (amd/skills, 398 stars), Agents (elevenlabs/skills, 481 stars), Local AI App Integration (amd/skills, 398 stars) and Codex Imagen (darkamenosa/codex-imagen, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Fish Audio API?

waker240 (a GitHub user) maintains it in waker240/FullVideoProductionSkill, which has 197 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on September 26, 2026.

Source: waker240/FullVideoProductionSkill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.