Agent skill

Listenhub Tts

by smallnest in smallnest/goal-workflow

使用 ListenHub API 将文本转换为语音(TTS)。支持三种模式:快速合成(/v1/tts)、 多角色脚本(/v1/speech)、长文本流式合成(/v1/flow-speech/episodes)。

MITAuto-check: notesMedia & Creative

Install Listenhub Tts

skills CLI
$ npx skills add smallnest/goal-workflow --skill listenhub-tts -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install smallnest/goal-workflow listenhub-tts --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/smallnest/goal-workflow.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/listenhub-tts .claude/skills/listenhub-tts && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
listenhub-tts
GitHub stars
291
Token cost
~1.5k tokens
SKILL.md length
354 words
Files
1
Skills in repo
20
Repo updated
First seen
Licence
MIT

At a glance

使用 ListenHub API 将文本转换为语音(TTS)。支持三种模式:快速合成(/v1/tts)、 多角色脚本(/v1/speech)、长文本流式合成(/v1/flow-speech/episodes)。

  • Works in 3 steps: 调用 GET /v1/speakers/list?language=zh… → 按 AskUserQuestion 展示音色列表供用户选择,格式如下 → 用户确认后使用选定的 speakerId
  • Tasks that involve Text to speech and voice
  • SKILL.md covers API 信息, 音色选择流程, 三种合成模式 and 音色列表查询, plus 5 more sections
  • Calls curl; reaches api.marswave.ai; needs LISTENHUB_API_KEY

What it does

Listenhub Tts is an agent skill from smallnest/goal-workflow. 使用 ListenHub API 将文本转换为语音(TTS)。支持三种模式:快速合成(/v1/tts)、 多角色脚本(/v1/speech)、长文本流式合成(/v1/flow-speech/episodes)。 音色未指定时自动获取音色列表供用户选择,默认使用 chat-girl-105-cn(晓曼)。 Use when user says: "tts", "text to speech", "语音合成", "文字转语音", "朗读", "生成语音", "生成音频", "转音频", "text to audio"

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice. The repository describes itself as: AI-driven development workflow with /prd, /goal, /review-it and /ship-it skills. The licence is MIT.

When your agent uses it

  • Tasks that involve Text to speech and voice

Example prompts

  • “text to speech”
  • “text to audio”
  • “/listenhub-tts”

Requirements

  • A credential in LISTENHUB_API_KEY
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Bash, AskUserQuestion

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. 调用 GET /v1/speakers/list?language=zh 获取可用音色列表
  2. 按 AskUserQuestion 展示音色列表供用户选择,格式如下
  3. 用户确认后使用选定的 speakerId

What it can do on your machine

Read from SKILL.md and the folder at commit b06ab3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Bash
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.marswave.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • LISTENHUB_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Listenhub Tts loads about 1.5k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 354 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Write, Edit, Bash, AskUserQuestion

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from smallnest/goal-workflow at commit b06ab3c, republished under its MIT licence (© smallnest). 354 words, ~1,532 tokens.

Download SKILL.mdSave it as .claude/skills/listenhub-tts/SKILL.md (or your agent's skills folder).
name
listenhub-tts
description
使用 ListenHub API 将文本转换为语音(TTS)。支持三种模式:快速合成(/v1/tts)、 多角色脚本(/v1/speech)、长文本流式合成(/v1/flow-speech/episodes)。 音色未指定时自动获取音色列表供用户选择,默认使用 chat-girl-105-cn(晓曼)。 Use when user says: "tts", "text to speech", "语音合成", "文字转语音", "朗读", "生成语音", "生成音频", "转音频", "text to audio"
allowed-tools
Read, Write, Edit, Bash, AskUserQuestion
user-invocable
true
metadata.trigger
TTS语音合成、text to speech、文字转语音、朗读文本、生成音频

ListenHub TTS: 文本转语音

使用 ListenHub OpenAPI 将文本转换为语音。支持三种合成模式,覆盖从短文本到长文本的全场景。

API 信息

  • Base URL: https://api.marswave.ai/openapi
  • 认证: Authorization: Bearer $LISTENHUB_API_KEY(从环境变量读取)
  • 前置检查: 调用任何 API 前先确认 LISTENHUB_API_KEY 环境变量已设置,未设置则提示用户配置

音色选择流程

用户已明确指定音色

直接使用用户指定的 speakerId,跳过选择流程。

用户未指定音色
  1. 调用 GET /v1/speakers/list?language=zh 获取可用音色列表
  2. 按 AskUserQuestion 展示音色列表供用户选择,格式如下:
    • 默认选中 chat-girl-105-cn(晓曼 dxqqq)
    • 列表展示:{name}({gender},{speakerId})
    • 附带每个音色的 demoAudioUrl 供参考
  3. 用户确认后使用选定的 speakerId
默认音色
字段值
speakerIdchat-girl-105-cn
名称晓曼 dxqqq

三种合成模式

模式一:快速合成(短文本,单音色)

适用场景: 短文本(< 1000 字),单音色,需要低延迟

接口: POST /v1/tts

请求体:

json
{
  "text": "要合成的文本",
  "speakerId": "chat-girl-105-cn",
  "format": "mp3",
  "sampleRate": 24000,
  "speed": 1.0
}
参数类型必填说明
textstring是要合成的文本
speakerIdstring是音色 ID
formatstring否输出格式,默认 mp3
sampleRateint否采样率,默认 24000
speedfloat否语速,默认 1.0,范围 0.5 ~ 2.0

响应: 直接返回 MP3 二进制流(Content-Type: audio/mpeg)

调用示例:

bash
curl -X POST "https://api.marswave.ai/openapi/v1/tts" \
  -H "Authorization: Bearer $LISTENHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text": "你好世界", "speakerId": "chat-girl-105-cn"}' \
  -o output.mp3
模式二:多角色脚本合成

适用场景: 多角色对话、播客、有声书片段,需要不同音色交替朗读

接口: POST /v1/speech

请求体:

json
{
  "script": [
    {
      "text": "你好,欢迎收听本期节目。",
      "speakerId": "chat-girl-105-cn"
    },
    {
      "text": "谢谢,今天我们来聊聊 AI。",
      "speakerId": "chat-boy-101-cn"
    }
  ],
  "format": "mp3",
  "sampleRate": 24000
}
参数类型必填说明
scriptarray是脚本数组,每项包含 text 和 speakerId
script[].textstring是该段文本
script[].speakerIdstring是该段的音色 ID
formatstring否输出格式,默认 mp3
sampleRateint否采样率,默认 24000

响应: JSON

json
{
  "audioUrl": "https://cdn.example.com/output.mp3",
  "duration": 12.5
}

调用示例:

bash
curl -X POST "https://api.marswave.ai/openapi/v1/speech" \
  -H "Authorization: Bearer $LISTENHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "script": [
      {"text": "你好,欢迎收听。", "speakerId": "chat-girl-105-cn"},
      {"text": "谢谢,我们开始吧。", "speakerId": "chat-boy-101-cn"}
    ]
  }'
模式三:长文本流式合成

适用场景: 长文本(> 1000 字),文章朗读,需要 AI 润色或分段处理

接口: POST /v1/flow-speech/episodes

请求体:

json
{
  "title": "文章标题",
  "content": "长文本内容...",
  "speakerId": "chat-girl-105-cn",
  "mode": "direct",
  "format": "mp3"
}
参数类型必填说明
titlestring是音频标题
contentstring否文本内容(与 contentUrl 二选一)
contentUrlstring否内容 URL(与 content 二选一)
speakerIdstring是音色 ID
modestring否direct(直接合成)或 aiPolish(AI 润色),默认 direct
formatstring否输出格式,默认 mp3

响应: JSON

json
{
  "episodeId": "ep_abc123",
  "status": "processing"
}

轮询获取结果:

bash
GET /v1/flow-speech/episodes/{episodeId}

轮询策略:

  1. 提交后等待 30 秒
  2. 之后每 10 秒轮询一次
  3. 直到 status 变为 completed 或 failed

轮询响应:

json
{
  "episodeId": "ep_abc123",
  "status": "completed",
  "audioUrl": "https://cdn.example.com/output.mp3",
  "duration": 180.5
}
status 值说明
processing合成中,继续轮询
completed合成完成,audioUrl 可用
failed合成失败,查看 errorMessage

调用示例:

bash
# 提交任务
curl -X POST "https://api.marswave.ai/openapi/v1/flow-speech/episodes" \
  -H "Authorization: Bearer $LISTENHUB_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "title": "AI 技术趋势",
    "content": "长文本内容...",
    "speakerId": "chat-girl-105-cn",
    "mode": "direct"
  }'

# 轮询结果
curl "https://api.marswave.ai/openapi/v1/flow-speech/episodes/ep_abc123" \
  -H "Authorization: Bearer $LISTENHUB_API_KEY"
Show full SKILL.md (133 more words)Show less

音色列表查询

接口: GET /v1/speakers/list

查询参数:

参数类型必填说明
languagestring否筛选语言,如 zh(中文)、en(英文)

响应:

json
{
  "speakers": [
    {
      "name": "晓曼 dxqqq",
      "speakerId": "chat-girl-105-cn",
      "demoAudioUrl": "https://cdn.example.com/demo.mp3",
      "gender": "female",
      "language": "zh"
    }
  ]
}

模式选择逻辑

根据用户输入自动选择最合适的模式:

条件模式
文本 ≤ 1000 字,单音色模式一:/v1/tts
多角色脚本,需要不同音色模式二:/v1/speech
文本 > 1000 字,或需要 AI 润色模式三:/v1/flow-speech/episodes
用户提供 URL 作为内容来源模式三:/v1/flow-speech/episodes

如果用户明确指定模式,优先使用用户指定的模式。

用户交互

音色选择

当用户未指定音色时,使用 AskUserQuestion 展示音色列表:

请选择音色(默认:晓曼 dxqqq):
A. 晓曼 dxqqq(女,chat-girl-105-cn)[默认]
B. [其他音色名称]([性别],[speakerId])
C. ...
合成参数

可选询问:

  • 语速 speed(默认 1.0)
  • 输出格式 format(默认 mp3)
  • 长文本模式:direct 还是 aiPolish(默认 direct)
  • 输出文件路径(默认 ./output.mp3)

输出

  1. 将音频保存到指定路径(默认 ./output.mp3)
  2. 输出合成摘要:
    • 使用的模式
    • 音色名称和 ID
    • 音频时长
    • 文件大小
    • 文件路径

错误处理

  • 401 Unauthorized: 提示用户检查 LISTENHUB_API_KEY 环境变量
  • 400 Bad Request: 检查请求参数,向用户报告具体错误
  • flow-speech failed: 报告 errorMessage,建议用户重试或切换模式
  • 网络错误: 提示检查网络连接,建议重试

完整示例

用户输入: "把这段文字转成语音:今天天气真好,适合出去散步。"

执行流程:

  1. 检查 LISTENHUB_API_KEY ✓
  2. 文本长度 < 1000 字,单音色 → 选择模式一 /v1/tts
  3. 用户未指定音色 → 默认使用 chat-girl-105-cn(晓曼)
  4. 调用 API 合成
  5. 保存到 ./output.mp3
  6. 输出摘要

用户输入: "用晓曼的声音朗读这篇文章:article.md"

执行流程:

  1. 读取 article.md 内容
  2. 检查文本长度 > 1000 字 → 选择模式三 /v1/flow-speech/episodes
  3. 音色已指定:chat-girl-105-cn(晓曼)
  4. 提交合成任务
  5. 轮询直到完成
  6. 下载音频保存到 ./article.mp3
  7. 输出摘要

© smallnest, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/listenhub-tts of smallnest/goal-workflow.

Open the folder on GitHubat commit b06ab3c

Compare with similar skills

Listenhub Tts next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Listenhub Tts compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Listenhub Tts this skillsmallnest/goal-workflow291—~1.5kAutomated safety check: NotesMIT
MoneyPrinterTurbo Video Generatorharry0703/MoneyPrinterTurbo129k—~2.1kAutomated safety check: WarnMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Openspec OnboardSAP/e-mobility-charging-stations-simulator22725 repos~3.5kAutomated safety check: PassMIT
Blog AudioAgriciDaniel/claude-blog2.3k1 repos~2.2kAutomated safety check: NotesMIT
Musictadaspetra/loop2962 repos~827Automated safety check: PassMIT

Similar skills

  • MoneyPrinterTurbo Video Generator

    harry0703/MoneyPrinterTurbo

    Installs and runs MoneyPrinterTurbo to turn a topic or script into a finished short video with voice-over, subtitles, stock footage and music.

    129k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Openspec Onboard

    SAP/e-mobility-charging-stations-simulator

    Official

    Guided onboarding for OpenSpec - walk through a complete workflow cycle with narration and real codebase work.

    227 GitHub starsUsed in 25 repos~3.5k tokens
    Media & CreativeAuto-check passed
  • Blog Audio

    AgriciDaniel/claude-blog

    Generate audio narration of blog posts using Google Gemini TTS.

    2.3k GitHub starsUsed in 1 repo~2.2k tokens
    Media & CreativeAuto-check: notes
  • Music

    tadaspetra/loop

    Generate music using ElevenLabs Music API. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~827 tokens
    Media & CreativeAuto-check passed
  • Create News Video

    hoquanghai/Auto-Create-Video

    Tạo video tin tức ngắn 9:16 (~60s) từ URL bài báo hoặc file .txt tiếng Việt.

    319 GitHub starsUsed in 1 repo~3.7k tokens
    Media & CreativeAuto-check passed

More from smallnest/goal-workflow

All 20 skills in this repo
  • Article Icons

    smallnest/goal-workflow

    Illustrate an article (Markdown, HTML, etc.) with animated-style icons from itshover.com/icons.

    291 GitHub stars~1.6k tokensUpdated 27 days ago
    Auto-check passed
  • Graph

    smallnest/goal-workflow

    Graph engineering for parallel task execution: convert a task, PRD, SPEC, or issue set into a dependency graph (DAG), layer it into supersteps, then implement each independent node concurrently with…

    291 GitHub stars~3.9k tokensUpdated 27 days ago
    Auto-check passed
  • Walkthrough

    smallnest/goal-workflow

    Generate a Phase-2 Walkthrough artifact (walkthrough.md) once implementation and verification are complete.

    291 GitHub stars~4.7k tokensUpdated 27 days ago
    Auto-check: notes
  • Insight Diagram

    smallnest/goal-workflow

    为任意项目生成 UML 图、架构图和流程图。分析代码库后让用户选择要生成的图表类型,使用 architecture-diagram skill 渲染为 HTML+SVG,保存到 docs/ 目录。适用于任何软件项目的文档可视化。

    291 GitHub stars~1.6k tokensUpdated 27 days ago
    Auto-check passed
  • Code To Spec

    smallnest/goal-workflow

    Reverse-engineer a SPEC document from an existing project. An agent skill from smallnest/goal-workflow.

    291 GitHub stars~2.7k tokensUpdated 27 days ago
    Auto-check passed
  • Design It

    smallnest/goal-workflow

    A skill your agent uses when turning a requirement, spec, or feature brief into a single self-contained HTML design document in a fixed house style — one styled HTML page with a table-of-contents…

    291 GitHub stars~1.1k tokensUpdated 27 days ago
    Auto-check passed

Questions about Listenhub Tts

What does Listenhub Tts do?

使用 ListenHub API 将文本转换为语音(TTS)。支持三种模式:快速合成(/v1/tts)、 多角色脚本(/v1/speech)、长文本流式合成(/v1/flow-speech/episodes)。. Listenhub Tts is an agent skill from smallnest/goal-workflow.

When should I use Listenhub Tts?

Listenhub Tts fits situations like: tasks that involve Text to speech and voice.

How do I install Listenhub Tts in Claude Code?

Run `npx skills add smallnest/goal-workflow --skill listenhub-tts -a claude-code`. Or copy the skill folder (skills/listenhub-tts in smallnest/goal-workflow) into .claude/skills/listenhub-tts in your project. Claude Code loads it when a task matches its description.

How do I install Listenhub Tts in Codex?

Run `npx skills add smallnest/goal-workflow --skill listenhub-tts -a codex`. Or copy the skill folder (skills/listenhub-tts in smallnest/goal-workflow) into .agents/skills/listenhub-tts in your project. Codex loads it when a task matches its description.

Can I use Listenhub Tts in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add smallnest/goal-workflow --skill listenhub-tts -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/listenhub-tts, .gemini/skills/listenhub-tts, .github/skills/listenhub-tts and .opencode/skills/listenhub-tts in your project.

What does Listenhub Tts need to run?

Going by SKILL.md and its folder, Listenhub Tts needs the command-line tools its instructions call (curl) and credentials named LISTENHUB_API_KEY. Our summary lists: A credential in LISTENHUB_API_KEY. Its frontmatter pre-approves these tools: Read, Write, Edit, Bash, AskUserQuestion.

Does Listenhub Tts access the network?

SKILL.md names 1 domain. In commands or code: api.marswave.ai; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Listenhub Tts safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Listenhub Tts use?

Listenhub Tts is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Listenhub Tts use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Listenhub Tts?

Skills that share tags, products or a category with Listenhub Tts: MoneyPrinterTurbo Video Generator (harry0703/MoneyPrinterTurbo, 129k stars), HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Openspec Onboard (SAP/e-mobility-charging-stations-simulator, 227 stars) and Blog Audio (AgriciDaniel/claude-blog, 2.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Listenhub Tts?

smallnest (a GitHub user) maintains it in smallnest/goal-workflow, which has 291 GitHub stars. The repository holds 20 skills in this directory. The repository was last updated on September 13, 2026.

Source: smallnest/goal-workflow on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.