Agent skill

Deepgram JS Text To Speech

by deepgram in deepgram/deepgram-js-sdk

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis.

MITAuto-check passedMedia & Creative

Install Deepgram JS Text To Speech

skills CLI
$ npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-text-to-speech -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install deepgram/deepgram-js-sdk deepgram-js-text-to-speech --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/deepgram/deepgram-js-sdk.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deepgram-js-text-to-speech .claude/skills/deepgram-js-text-to-speech && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepgram-js-text-to-speech
GitHub stars
276
Token cost
~1.4k tokens
SKILL.md length
393 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis.

  • Works in 5 steps: In-repo reference: reference.md → Speak… → Canonical OpenAPI (REST):… → Canonical AsyncAPI (WSS):… → …
  • Reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis
  • SKILL.md covers When to use this product, Authentication, Quick start — REST (one-shot) and Quick start — WebSocket…, plus 6 more sections
  • Calls npx; needs DEEPGRAM_API_KEY

What it does

Deepgram JS Text To Speech is an agent skill from deepgram/deepgram-js-sdk. Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis. Covers one-shot REST via client.speak.v1.audio.generate and streaming WebSocket via client.speak.v1.createConnection() / connect(). Use deepgram-js-voice-agent when you need full-duplex STT + LLM + TTS instead of one-way synthesis. Triggers include "TTS", "text to speech", "speak", "aura", "streaming TTS", and "speak.v1".

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Text to speech and voice and Speech recognition and synthesis. It works with Deepgram, JavaScript and TypeScript. The repository describes itself as: Official JavaScript SDK for Deepgram. The licence is MIT.

When your agent uses it

  • Reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis
  • Tasks that involve Text to speech and voice
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “text to speech”
  • “streaming TTS”
  • “speak.v1”
  • “/deepgram-js-text-to-speech”

Requirements

  • Python 3
  • Node.js
  • A credential in DEEPGRAM_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. In-repo reference: reference.md → Speak V1 Audio for REST; WSS behavior lives in src/CustomClient.ts and…
  2. Canonical OpenAPI (REST): https://developers.deepgram.com/openapi.yaml
  3. Canonical AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  4. Context7: library ID /llmstxt/developers_deepgram_llms_txt
  5. Product docs

What it can do on your machine

Read from SKILL.md and the folder at commit 108a127. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.deepgram.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEEPGRAM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deepgram JS Text To Speech loads about 1.4k tokens when it runs. Until then it costs about 124 tokens; SKILL.md has 393 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~124
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from deepgram/deepgram-js-sdk at commit 108a127, republished under its MIT licence (© deepgram). 393 words, ~1,390 tokens.

Download SKILL.mdSave it as .claude/skills/deepgram-js-text-to-speech/SKILL.md (or your agent's skills folder).
name
deepgram-js-text-to-speech
description
Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (`/v1/speak`) for audio synthesis. Covers one-shot REST via `client.speak.v1.audio.generate` and streaming WebSocket via `client.speak.v1.createConnection()` / `connect()`. Use `deepgram-js-voice-agent` when you need full-duplex STT + LLM + TTS instead of one-way synthesis. Triggers include "TTS", "text to speech", "speak", "aura", "streaming TTS", and "speak.v1".

Using Deepgram Text-to-Speech (JavaScript / TypeScript SDK)

Convert text to audio with one-shot REST generation or low-latency streaming synthesis via /v1/speak.

When to use this product

  • REST (client.speak.v1.audio.generate) — render finished text into an audio response. Best for downloadable files, pre-generated prompts, batch synthesis.
  • WebSocket (client.speak.v1.createConnection() / connect()) — stream text in and receive audio out with lower latency. Best when an LLM is still producing tokens.

Use a different skill when:

  • You need the agent to also listen, think, and handle barge-in → deepgram-js-voice-agent.

Authentication

js
require("dotenv").config();

const { DeepgramClient } = require("@deepgram/sdk");

const deepgramClient = new DeepgramClient({
  apiKey: process.env.DEEPGRAM_API_KEY,
});

The repo examples use require("../dist/cjs/index.js"), but application code should normally import from @deepgram/sdk.

Quick start — REST (one-shot)

From examples/10-text-to-speech-single.ts:

js
const data = await deepgramClient.speak.v1.audio.generate({
  text: "Hello, this is a test of Deepgram's text-to-speech API.",
  model: "aura-2-thalia-en",
  encoding: "linear16",
  container: "wav",
});

console.log("Audio generated successfully", data);

generate(...) returns a BinaryResponse, not JSON. See examples/25-binary-response.ts for .stream(), .arrayBuffer(), .blob(), and .bytes() handling.

Quick start — WebSocket (streaming)

From examples/11-text-to-speech-streaming.ts:

js
const deepgramConnection = await deepgramClient.speak.v1.createConnection({
  model: "aura-2-thalia-en",
  encoding: "linear16",
});

deepgramConnection.on("message", (data) => {
  if (typeof data === "string" || data instanceof ArrayBuffer || data instanceof Blob) {
    console.log("Audio received");
  } else if (data.type === "Flushed") {
    deepgramConnection.close();
  }
});

deepgramConnection.connect();
await deepgramConnection.waitForOpen();

deepgramConnection.sendText({ type: "Speak", text: "Hello from streaming TTS." });
deepgramConnection.sendFlush({ type: "Flush" });

Key parameters / API surface

  • REST & WSS: model, encoding, sample_rate, container, bit_rate, callback, callback_method, tag, mip_opt_out.
  • REST response surface (examples/25-binary-response.ts): response.stream(), response.arrayBuffer(), response.blob(), response.bytes(), response.bodyUsed.
  • WSS client messages (src/api/resources/speak/resources/v1/client/Socket.ts): sendText(...), sendFlush(...), sendClear(...), sendClose(...).
  • WSS server events: binary audio payloads plus Metadata, Flushed, Cleared, Warning.

Limitations

Unlike the Python SDK, this repo does not include a hand-written TextBuilder helper. If you want incremental token buffering before sendText(...), build that helper in your application layer.

API reference (layered)

  1. In-repo reference: reference.md → Speak V1 Audio for REST; WSS behavior lives in src/CustomClient.ts and src/api/resources/speak/resources/v1/client/{Client,Socket}.ts.
  2. Canonical OpenAPI (REST): https://developers.deepgram.com/openapi.yaml
  3. Canonical AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  4. Context7: library ID /llmstxt/developers_deepgram_llms_txt
  5. Product docs:
Show full SKILL.md (160 more words)Show less

Gotchas

  1. REST returns binary, not JSON. Treat the result like a streamed/binary body.
  2. Use the custom client wrapper. src/CustomClient.ts patches binary WebSocket handling; the generated socket assumes JSON too aggressively.
  3. createConnection() is lazy. Register handlers, then call connect() and waitForOpen().
  4. Send Flush after your text. Without sendFlush({ type: "Flush" }), trailing audio may not be emitted promptly.
  5. Streaming text is structured JSON. Send { type: "Speak", text }, not a raw string.
  6. Audio payload shape varies by runtime. The same handler may receive string, ArrayBuffer, or Blob.
  7. Pick encoding/container/sample rate that match your sink. Mismatches show up as static, silence, or unplayable files.

Example files in this repo

  • examples/10-text-to-speech-single.ts
  • examples/11-text-to-speech-streaming.ts
  • examples/25-binary-response.ts

Central product skills

For cross-language Deepgram product knowledge — the consolidated API reference, documentation finder, focused runnable recipes, third-party integration examples, and MCP setup — install the central skills:

bash
npx skills add deepgram/skills

This SDK ships language-idiomatic code skills; deepgram/skills ships cross-language product knowledge (see api, docs, recipes, examples, starters, setup-mcp).

© deepgram, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/deepgram-js-text-to-speech of deepgram/deepgram-js-sdk.

Open the folder on GitHubat commit 108a127

Compare with similar skills

Deepgram JS Text To Speech next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepgram JS Text To Speech compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepgram JS Text To Speech this skilldeepgram/deepgram-js-sdk276—~1.4kAutomated safety check: PassMIT
Speech To Texttadaspetra/loop2962 repos~2kAutomated safety check: PassMIT
Deepgram Python Text-to-Speechdeepgram/deepgram-python-sdk469—~1.8kAutomated safety check: PassMIT
Speech Engineelevenlabs/skills482—~2.5kAutomated safety check: WarnMIT
Voice AI Developmentdavila7/claude-code-templates33k5 repos~2.1kAutomated safety check: PassMIT
Azure AImicrosoft/GitHub-Copilot-for-Azure2551 repos~852Automated safety check: PassMIT

Similar skills

  • Speech To Text

    tadaspetra/loop

    Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~2k tokens
    Media & CreativeAuto-check passed
  • Deepgram Python Text-to-Speech

    deepgram/deepgram-python-sdk

    Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper.

    469 GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated today
    Media & CreativeAuto-check: warnings
  • Voice AI Development

    davila7/claude-code-templates

    Expert in building voice AI applications - from real-time voice agents to voice-enabled apps.

    33k GitHub starsUsed in 5 repos~2.1k tokens
    Media & CreativeAuto-check passed
  • Azure AI

    microsoft/GitHub-Copilot-for-Azure

    Official

    A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.

    255 GitHub starsUsed in 1 repo~852 tokens
    Media & CreativeAuto-check passed
  • Deepgram

    Anil-matcha/awesome-muse-connectors

    Deepgram speech AI: transcribe audio to text and synthesize speech (TTS).

    1.3k GitHub stars~857 tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

More from deepgram/deepgram-js-sdk

  • Deepgram JS Audio Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…

    276 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Deepgram JS Conversational Stt

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription.

    276 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Deepgram JS Management API

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Management APIs for projects, API keys, members, invites, requests, usage, billing, models…

    276 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Deepgram JS Speech To Text

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription.

    276 GitHub stars~1.8k tokensUpdated today
    Auto-check passed
  • Deepgram JS Text Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text Intelligence / Read (/v1/read) for sentiment, summarization, topic detection, and intent…

    276 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Deepgram JS Voice Agent

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that builds an interactive voice agent via agent.deepgram.com/v1/agent/converse.

    276 GitHub stars~1.6k tokensUpdated today
    Auto-check passed

Questions about Deepgram JS Text To Speech

What does Deepgram JS Text To Speech do?

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis. Deepgram JS Text To Speech is an agent skill from deepgram/deepgram-js-sdk. Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis.

When should I use Deepgram JS Text To Speech?

Deepgram JS Text To Speech fits situations like: reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis; tasks that involve Text to speech and voice; tasks that involve Speech recognition and synthesis.

How do I install Deepgram JS Text To Speech in Claude Code?

Run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-text-to-speech -a claude-code`. Or copy the skill folder (.agents/skills/deepgram-js-text-to-speech in deepgram/deepgram-js-sdk) into .claude/skills/deepgram-js-text-to-speech in your project. Claude Code loads it when a task matches its description.

How do I install Deepgram JS Text To Speech in Codex?

Run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-text-to-speech -a codex`. Or copy the skill folder (.agents/skills/deepgram-js-text-to-speech in deepgram/deepgram-js-sdk) into .agents/skills/deepgram-js-text-to-speech in your project. Codex loads it when a task matches its description.

Can I use Deepgram JS Text To Speech in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-text-to-speech -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepgram-js-text-to-speech, .gemini/skills/deepgram-js-text-to-speech, .github/skills/deepgram-js-text-to-speech and .opencode/skills/deepgram-js-text-to-speech in your project.

What does Deepgram JS Text To Speech need to run?

Going by SKILL.md and its folder, Deepgram JS Text To Speech needs the command-line tools its instructions call (npx) and credentials named DEEPGRAM_API_KEY. Our summary lists: Python 3; Node.js; A credential in DEEPGRAM_API_KEY.

Does Deepgram JS Text To Speech access the network?

SKILL.md names 1 domain. As links in the text: developers.deepgram.com. This is read from the text; nothing was executed.

Is Deepgram JS Text To Speech safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deepgram JS Text To Speech use?

Deepgram JS Text To Speech is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepgram JS Text To Speech use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deepgram JS Text To Speech?

Skills that share tags, products or a category with Deepgram JS Text To Speech: Speech To Text (tadaspetra/loop, 296 stars), Deepgram Python Text-to-Speech (deepgram/deepgram-python-sdk, 469 stars), Speech Engine (elevenlabs/skills, 482 stars) and Voice AI Development (davila7/claude-code-templates, 33k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepgram JS Text To Speech?

deepgram (a GitHub organization) maintains it in deepgram/deepgram-js-sdk, which has 276 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: deepgram/deepgram-js-sdk on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.