Agent skill

Deepgram JS Conversational Stt

by deepgram in deepgram/deepgram-js-sdk

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription.

MITAuto-check passedAI & LLM Engineering

Install Deepgram JS Conversational Stt

skills CLI
$ npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-conversational-stt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install deepgram/deepgram-js-sdk deepgram-js-conversational-stt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/deepgram/deepgram-js-sdk.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deepgram-js-conversational-stt .claude/skills/deepgram-js-conversational-stt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepgram-js-conversational-stt
GitHub stars
276
Token cost
~1.3k tokens
SKILL.md length
364 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription.

  • Works in 5 steps: In-repo reference: reference.md does not… → Canonical OpenAPI (REST):… → Canonical AsyncAPI (WSS):… → …
  • Conversational STT
  • SKILL.md covers When to use this product, Authentication, Quick start and Key parameters / API surface, plus 5 more sections
  • Calls npx; needs DEEPGRAM_API_KEY

What it does

Deepgram JS Conversational Stt is an agent skill from deepgram/deepgram-js-sdk. Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription. Covers client.listen.v2.createConnection() / connect(), Flux models, and turn events like TurnInfo. Use deepgram-js-speech-to-text for standard v1 ASR and deepgram-js-voice-agent for full-duplex assistants. Triggers include "flux", "v2 listen", "conversational STT", "turn detection", "end of turn", "EOT", and "listen.v2".

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Speech recognition and synthesis and Transcription. It works with Deepgram, JavaScript and TypeScript. The repository describes itself as: Official JavaScript SDK for Deepgram. The licence is MIT.

When your agent uses it

  • Conversational STT
  • Tasks that involve Speech recognition and synthesis
  • Tasks that involve Transcription

Example prompts

  • “v2 listen”
  • “conversational STT”
  • “turn detection”
  • “/deepgram-js-conversational-stt”

Requirements

  • Node.js
  • A credential in DEEPGRAM_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. In-repo reference: reference.md does not currently document listen.v2; use src/CustomClient.ts and…
  2. Canonical OpenAPI (REST): https://developers.deepgram.com/openapi.yaml
  3. Canonical AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  4. Context7: library ID /llmstxt/developers_deepgram_llms_txt
  5. Product docs

What it can do on your machine

Read from SKILL.md and the folder at commit 108a127. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.deepgram.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEEPGRAM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deepgram JS Conversational Stt loads about 1.3k tokens when it runs. Until then it costs about 133 tokens; SKILL.md has 364 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~133
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from deepgram/deepgram-js-sdk at commit 108a127, republished under its MIT licence (© deepgram). 364 words, ~1,312 tokens.

Download SKILL.mdSave it as .claude/skills/deepgram-js-conversational-stt/SKILL.md (or your agent's skills folder).
name
deepgram-js-conversational-stt
description
Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (`/v2/listen`) for turn-aware streaming transcription. Covers `client.listen.v2.createConnection()` / `connect()`, Flux models, and turn events like `TurnInfo`. Use `deepgram-js-speech-to-text` for standard v1 ASR and `deepgram-js-voice-agent` for full-duplex assistants. Triggers include "flux", "v2 listen", "conversational STT", "turn detection", "end of turn", "EOT", and "listen.v2".

Using Deepgram Conversational STT / Flux (JavaScript / TypeScript SDK)

Turn-aware streaming STT via /v2/listen for conversational audio and explicit turn events.

When to use this product

  • You need turn-aware transcription, not just a word stream.
  • You want Flux-style events like TurnInfo and Connected.
  • You are building a conversational interface but do not want the full Voice Agent runtime.

Use a different skill when:

  • You want general v1 transcription or prerecorded REST → deepgram-js-speech-to-text.
  • You want a hosted assistant with think + speak built in → deepgram-js-voice-agent.
  • You want analytics overlays like sentiment and summaries → deepgram-js-audio-intelligence.

Authentication

js
require("dotenv").config();

const { DeepgramClient } = require("@deepgram/sdk");

const deepgramClient = new DeepgramClient({
  apiKey: process.env.DEEPGRAM_API_KEY,
});

Quick start

From examples/26-transcription-live-websocket-v2.ts:

js
const deepgramConnection = await deepgramClient.listen.v2.createConnection({
  model: "flux-general-en",
});

deepgramConnection.on("message", (data) => {
  if (data.type === "Connected") {
    console.log("Connected:", data);
  } else if (data.type === "TurnInfo") {
    console.log("Turn Info:", data);
  } else if (data.type === "FatalError") {
    deepgramConnection.close();
  }
});

deepgramConnection.connect();
await deepgramConnection.waitForOpen();

// Swap this for a live mic capture in real apps; the repo example uses
// `createReadStream` over a sample file.
const { createReadStream } = require("node:fs");
const audioStream = createReadStream("samples/spacewalk.wav");

audioStream.on("data", (chunk) => {
  deepgramConnection.sendMedia(chunk);
});

audioStream.on("end", () => {
  deepgramConnection.sendCloseStream({ type: "CloseStream" });
});

Key parameters / API surface

  • Connect args from src/api/resources/listen/resources/v2/client/Client.ts: model, encoding, sample_rate, eager_eot_threshold, eot_threshold, eot_timeout_ms, keyterm, tag, mip_opt_out.
  • Socket methods from src/api/resources/listen/resources/v2/client/Socket.ts: sendMedia(...), sendCloseStream(...), sendListenV2Configure(...), waitForOpen().
  • Custom wrapper additions from src/CustomClient.ts: createConnection(...) alias and Node-only ping(...) helper.

Limitations

  • The current generated ListenV2Model type only exposes "flux-general-en".
  • Product docs mention broader Flux capabilities, but flux-general-multi / language_hint are not surfaced in this repo's generated types today.

API reference (layered)

  1. In-repo reference: reference.md does not currently document listen.v2; use src/CustomClient.ts and src/api/resources/listen/resources/v2/client/{Client,Socket}.ts as the in-repo source of truth.
  2. Canonical OpenAPI (REST): https://developers.deepgram.com/openapi.yaml
  3. Canonical AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  4. Context7: library ID /llmstxt/developers_deepgram_llms_txt
  5. Product docs:
Show full SKILL.md (162 more words)Show less

Gotchas

  1. There is no REST path here. This skill is /v2/listen WebSocket only.
  2. Current repo typing only allows flux-general-en. Treat multilingual Flux support as a product capability not yet reflected in this SDK surface.
  3. Close with sendCloseStream, not sendFinalize. Finalize is the v1 pattern.
  4. ping() is Node-only. src/CustomClient.ts throws in browsers because browser WS ping frames are not user-exposed.
  5. createConnection() is lazy. Call connect() after registering handlers.
  6. The example explicitly warns about access/availability. A 400 may mean your account or endpoint is not enabled yet, not that your socket code is wrong.
  7. Omit encoding for containerized audio. Keep it for raw PCM/Opus streams.

Example files in this repo

  • examples/26-transcription-live-websocket-v2.ts
  • examples/27-deepgram-session-header.ts

Central product skills

For cross-language Deepgram product knowledge — the consolidated API reference, documentation finder, focused runnable recipes, third-party integration examples, and MCP setup — install the central skills:

bash
npx skills add deepgram/skills

This SDK ships language-idiomatic code skills; deepgram/skills ships cross-language product knowledge (see api, docs, recipes, examples, starters, setup-mcp).

© deepgram, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/deepgram-js-conversational-stt of deepgram/deepgram-js-sdk.

Open the folder on GitHubat commit 108a127

Compare with similar skills

Deepgram JS Conversational Stt next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepgram JS Conversational Stt compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepgram JS Conversational Stt this skilldeepgram/deepgram-js-sdk276—~1.3kAutomated safety check: PassMIT
Deepgram Audio Intelligence for Pythondeepgram/deepgram-python-sdk469—~2.3kAutomated safety check: PassMIT
Transcribe Anythingswyxio/skills176—~8.5kAutomated safety check: PassMIT
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
Deepgram Flux Conversational STTdeepgram/deepgram-python-sdk469—~1.8kAutomated safety check: PassMIT

Similar skills

  • Deepgram Audio Intelligence for Python

    deepgram/deepgram-python-sdk

    Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python.

    469 GitHub stars~2.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Transcribe Anything

    swyxio/skills

    Transcribes audio and video files to text using pluggable ASR backends.

    176 GitHub stars~8.5k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated 3 days ago
    Backend & APIsAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Deepgram Flux Conversational STT

    deepgram/deepgram-python-sdk

    Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection.

    469 GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Transformers.js

    huggingface/skills

    Official

    Runs pre-trained Hugging Face models in JavaScript or TypeScript with Transformers.js, in browsers or Node.js, Bun and Deno, for text, vision, audio and multimodal tasks.

    11k GitHub starsUsed in 1 repo~6.2k tokens
    AI & LLM EngineeringAuto-check passed

More from deepgram/deepgram-js-sdk

  • Deepgram JS Audio Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…

    276 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Management API

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Management APIs for projects, API keys, members, invites, requests, usage, billing, models…

    276 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Speech To Text

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription.

    276 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Text Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text Intelligence / Read (/v1/read) for sentiment, summarization, topic detection, and intent…

    276 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Text To Speech

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis.

    276 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Voice Agent

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that builds an interactive voice agent via agent.deepgram.com/v1/agent/converse.

    276 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed

Questions about Deepgram JS Conversational Stt

What does Deepgram JS Conversational Stt do?

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription. Deepgram JS Conversational Stt is an agent skill from deepgram/deepgram-js-sdk. Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription.

When should I use Deepgram JS Conversational Stt?

Deepgram JS Conversational Stt fits situations like: conversational STT; tasks that involve Speech recognition and synthesis; tasks that involve Transcription.

How do I install Deepgram JS Conversational Stt in Claude Code?

Run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-conversational-stt -a claude-code`. Or copy the skill folder (.agents/skills/deepgram-js-conversational-stt in deepgram/deepgram-js-sdk) into .claude/skills/deepgram-js-conversational-stt in your project. Claude Code loads it when a task matches its description.

How do I install Deepgram JS Conversational Stt in Codex?

Run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-conversational-stt -a codex`. Or copy the skill folder (.agents/skills/deepgram-js-conversational-stt in deepgram/deepgram-js-sdk) into .agents/skills/deepgram-js-conversational-stt in your project. Codex loads it when a task matches its description.

Can I use Deepgram JS Conversational Stt in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-conversational-stt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepgram-js-conversational-stt, .gemini/skills/deepgram-js-conversational-stt, .github/skills/deepgram-js-conversational-stt and .opencode/skills/deepgram-js-conversational-stt in your project.

What does Deepgram JS Conversational Stt need to run?

Going by SKILL.md and its folder, Deepgram JS Conversational Stt needs the command-line tools its instructions call (npx) and credentials named DEEPGRAM_API_KEY. Our summary lists: Node.js; A credential in DEEPGRAM_API_KEY.

Does Deepgram JS Conversational Stt access the network?

SKILL.md names 1 domain. As links in the text: developers.deepgram.com. This is read from the text; nothing was executed.

Is Deepgram JS Conversational Stt safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deepgram JS Conversational Stt use?

Deepgram JS Conversational Stt is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepgram JS Conversational Stt use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deepgram JS Conversational Stt?

Skills that share tags, products or a category with Deepgram JS Conversational Stt: Deepgram Audio Intelligence for Python (deepgram/deepgram-python-sdk, 469 stars), Transcribe Anything (swyxio/skills, 176 stars), Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars) and 9Router Speech-to-Text (decolua/9router, 31k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepgram JS Conversational Stt?

deepgram (a GitHub organization) maintains it in deepgram/deepgram-js-sdk, which has 276 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: deepgram/deepgram-js-sdk on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.