Agent skill

Deepgram JS Speech To Text

by deepgram in deepgram/deepgram-js-sdk

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription.

MITAuto-check passedMedia & Creative

Install Deepgram JS Speech To Text

skills CLI
$ npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-speech-to-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install deepgram/deepgram-js-sdk deepgram-js-speech-to-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/deepgram/deepgram-js-sdk.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deepgram-js-speech-to-text .claude/skills/deepgram-js-speech-to-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepgram-js-speech-to-text
GitHub stars
276
Token cost
~1.8k tokens
SKILL.md length
459 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription.

  • Works in 5 steps: In-repo reference: reference.md → Listen… → Canonical OpenAPI (REST):… → Canonical AsyncAPI (WSS):… → …
  • Reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded
  • SKILL.md covers When to use this product, Authentication, Quick start — REST… and Quick start — REST…, plus 6 more sections
  • Calls npx; reaches dpgr.am; needs DEEPGRAM_API_KEY

What it does

Deepgram JS Speech To Text is an agent skill from deepgram/deepgram-js-sdk. Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription. Covers client.listen.v1.media.transcribeUrl / transcribeFile (REST) plus client.listen.v1.createConnection() / connect() (WebSocket). Use deepgram-js-audio-intelligence for summarize/sentiment/topics/diarize overlays, deepgram-js-conversational-stt for Flux turn-taking on /v2/listen, and deepgram-js-voice-agent for full-duplex assistants. Triggers include…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with Deepgram, JavaScript and TypeScript. The repository describes itself as: Official JavaScript SDK for Deepgram. The licence is MIT.

When your agent uses it

  • Reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded
  • Live audio transcription
  • Include transcribe
  • Live transcription

Example prompts

  • “transcribe”
  • “speech to text”
  • “listen.v1”
  • “/deepgram-js-speech-to-text”

Requirements

  • Node.js
  • A credential in DEEPGRAM_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. In-repo reference: reference.md → Listen V1 Media for REST; WSS behavior lives in src/CustomClient.ts and…
  2. Canonical OpenAPI (REST): https://developers.deepgram.com/openapi.yaml
  3. Canonical AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  4. Context7: library ID /llmstxt/developers_deepgram_llms_txt
  5. Product docs

What it can do on your machine

Read from SKILL.md and the folder at commit 108a127. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dpgr.am

    Also links to:

    • developers.deepgram.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEEPGRAM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deepgram JS Speech To Text loads about 1.8k tokens when it runs. Until then it costs about 170 tokens; SKILL.md has 459 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~170
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from deepgram/deepgram-js-sdk at commit 108a127, republished under its MIT licence (© deepgram). 459 words, ~1,804 tokens.

Download SKILL.mdSave it as .claude/skills/deepgram-js-speech-to-text/SKILL.md (or your agent's skills folder).
name
deepgram-js-speech-to-text
description
Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (`/v1/listen`) for prerecorded or live audio transcription. Covers `client.listen.v1.media.transcribeUrl` / `transcribeFile` (REST) plus `client.listen.v1.createConnection()` / `connect()` (WebSocket). Use `deepgram-js-audio-intelligence` for summarize/sentiment/topics/diarize overlays, `deepgram-js-conversational-stt` for Flux turn-taking on `/v2/listen`, and `deepgram-js-voice-agent` for full-duplex assistants. Triggers include "transcribe", "speech to text", "STT", "listen.v1", "nova-3", "live transcription", and "websocket transcription".

Using Deepgram Speech-to-Text (JavaScript / TypeScript SDK)

Basic transcription for prerecorded audio (REST) or live audio (WebSocket) via /v1/listen.

When to use this product

  • REST (client.listen.v1.media.transcribeUrl / transcribeFile) — one-shot transcription of a finished URL or file. Good for batch jobs, caption generation, offline processing.
  • WebSocket (client.listen.v1.createConnection() / connect()) — continuous streaming transcription. Good for live captions, microphone audio, telephony streams, browser or Node realtime apps.

Use a different skill when:

  • You also want summaries, topics, intents, sentiment, language detection, or redaction guidance on the same /v1/listen call → deepgram-js-audio-intelligence.
  • You need Flux turn-taking and end-of-turn events on /v2/listen → deepgram-js-conversational-stt.
  • You need a full interactive assistant with STT + LLM + TTS over one socket → deepgram-js-voice-agent.

Authentication

js
require("dotenv").config();

const { DeepgramClient } = require("@deepgram/sdk");

const deepgramClient = new DeepgramClient({
  apiKey: process.env.DEEPGRAM_API_KEY,
});

Use the exported DeepgramClient from src/CustomClient.ts, not DefaultDeepgramClient. The wrapper adds the required Token auth prefix, session headers, and patched WebSocket behavior.

Quick start — REST (prerecorded URL)

From examples/04-transcription-prerecorded-url.ts:

js
const data = await deepgramClient.listen.v1.media.transcribeUrl({
  url: "https://dpgr.am/spacewalk.wav",
  model: "nova-3",
  language: "en",
  punctuate: true,
  paragraphs: true,
  utterances: true,
});

console.log(
  "Transcription:",
  data.results?.channels?.[0]?.alternatives?.[0]?.transcript,
);

Quick start — REST (prerecorded file)

From examples/05-transcription-prerecorded-file.ts:

js
const { createReadStream } = require("fs");

const data = await deepgramClient.listen.v1.media.transcribeFile(
  createReadStream("./examples/spacewalk.wav"),
  {
    model: "nova-3",
    language: "en",
    punctuate: true,
    paragraphs: true,
    utterances: true,
    smart_format: true,
  }
);

transcribeFile(...) accepts multiple upload shapes in this SDK: fs.ReadStream, Buffer, ReadableStream, Blob, File, ArrayBuffer, and Uint8Array (see examples/23-file-upload-types.ts).

Quick start — WebSocket (live streaming)

From examples/07-transcription-live-websocket.ts:

js
const deepgramConnection = await deepgramClient.listen.v1.createConnection({
  model: "nova-3",
  language: "en",
  punctuate: "true",
  interim_results: "true",
});

deepgramConnection.on("message", (data) => {
  if (data.type === "Results") {
    console.log("Transcript:", data);
  }
});

deepgramConnection.connect();
await deepgramConnection.waitForOpen();

// Swap this for a mic capture (e.g. `node-microphone` / `MediaRecorder`)
// in real apps; the repo examples use `createReadStream` over a sample WAV.
const { createReadStream } = require("node:fs");
const audioStream = createReadStream("samples/spacewalk.wav");

audioStream.on("data", (chunk) => {
  deepgramConnection.sendMedia(chunk);
});

audioStream.on("end", () => {
  deepgramConnection.sendFinalize({ type: "Finalize" });
});

The repo examples use the two-step socket flow: createConnection() → register handlers → connect() → waitForOpen().

Key parameters / API surface

  • REST: model, language, punctuate, smart_format, paragraphs, utterances, multichannel, numerals, search, keyterm, keywords, encoding, sample_rate, callback, tag.
  • WSS connect args (src/api/resources/listen/resources/v1/client/Client.ts): model is required; common realtime flags include language, interim_results, endpointing, utterance_end_ms, vad_events, encoding, sample_rate, multichannel, punctuate, smart_format.
  • WSS client messages (src/api/resources/listen/resources/v1/client/Socket.ts): sendMedia(...), sendFinalize(...), sendCloseStream(...), sendKeepAlive(...).
  • WSS server events: Results, Metadata, UtteranceEnd, SpeechStarted.

API reference (layered)

  1. In-repo reference: reference.md → Listen V1 Media for REST; WSS behavior lives in src/CustomClient.ts and src/api/resources/listen/resources/v1/client/{Client,Socket}.ts.
  2. Canonical OpenAPI (REST): https://developers.deepgram.com/openapi.yaml
  3. Canonical AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  4. Context7: library ID /llmstxt/developers_deepgram_llms_txt
  5. Product docs:
Show full SKILL.md (181 more words)Show less

Gotchas

  1. Use DeepgramClient, not DefaultDeepgramClient. The custom wrapper adds Token auth, session IDs, browser WS auth protocols, and patched sockets.
  2. Repo examples are two-stage for WSS. createConnection() does not open the socket; call connect() and usually waitForOpen().
  3. Finalize before closing v1 streams. sendFinalize({ type: "Finalize" }) flushes the final partial.
  4. Keep idle streams alive. Use audio or sendKeepAlive({ type: "KeepAlive" }) on long pauses.
  5. Raw audio metadata must match reality. If you send PCM, encoding and sample_rate must match the bytes.
  6. Browser auth differs from Node auth. In browsers, the wrapper moves auth/session info into WebSocket subprotocols because custom headers are unavailable.
  7. Use /v2/listen only for Flux. If you need turn-aware conversational STT, switch skills instead of forcing v1.

Example files in this repo

  • examples/04-transcription-prerecorded-url.ts
  • examples/05-transcription-prerecorded-file.ts
  • examples/06-transcription-prerecorded-callback.ts
  • examples/07-transcription-live-websocket.ts
  • examples/08-transcription-captions.ts
  • examples/23-file-upload-types.ts
  • examples/27-deepgram-session-header.ts

Central product skills

For cross-language Deepgram product knowledge — the consolidated API reference, documentation finder, focused runnable recipes, third-party integration examples, and MCP setup — install the central skills:

bash
npx skills add deepgram/skills

This SDK ships language-idiomatic code skills; deepgram/skills ships cross-language product knowledge (see api, docs, recipes, examples, starters, setup-mcp).

© deepgram, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/deepgram-js-speech-to-text of deepgram/deepgram-js-sdk.

Open the folder on GitHubat commit 108a127

Compare with similar skills

Deepgram JS Speech To Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepgram JS Speech To Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepgram JS Speech To Text this skilldeepgram/deepgram-js-sdk276—~1.8kAutomated safety check: PassMIT
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0
9Router Speech-to-Textdecolua/9router31k—~914Automated safety check: PassMIT
Speech Engineelevenlabs/skills482—~2.5kAutomated safety check: WarnMIT
Azure AImicrosoft/GitHub-Copilot-for-Azure2551 repos~852Automated safety check: PassMIT
Deepgram Core Workflow Bjeremylongshore/tons-of-skills-marketplace2.8k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Speech Engine

    elevenlabs/skills

    Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.

    482 GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: warnings
  • Azure AI

    microsoft/GitHub-Copilot-for-Azure

    Official

    A skill your agent uses for Azure AI: Search, Speech, OpenAI, Document Intelligence.

    255 GitHub starsUsed in 1 repo~852 tokens
    Media & CreativeAuto-check passed
  • Deepgram Core Workflow B

    jeremylongshore/tons-of-skills-marketplace

    Implement real-time streaming transcription with Deepgram WebSocket.

    2.8k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Speech To Text

    tadaspetra/loop

    Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~2k tokens
    Media & CreativeAuto-check passed

More from deepgram/deepgram-js-sdk

  • Deepgram JS Audio Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…

    276 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Conversational Stt

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription.

    276 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Management API

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Management APIs for projects, API keys, members, invites, requests, usage, billing, models…

    276 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Text Intelligence

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text Intelligence / Read (/v1/read) for sentiment, summarization, topic detection, and intent…

    276 GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Text To Speech

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Text-to-Speech v1 (/v1/speak) for audio synthesis.

    276 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Deepgram JS Voice Agent

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that builds an interactive voice agent via agent.deepgram.com/v1/agent/converse.

    276 GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed

Questions about Deepgram JS Speech To Text

What does Deepgram JS Speech To Text do?

A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription. Deepgram JS Speech To Text is an agent skill from deepgram/deepgram-js-sdk. Use when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded or live audio transcription.

When should I use Deepgram JS Speech To Text?

Deepgram JS Speech To Text fits situations like: reviewing JavaScript/TypeScript in this repo that calls Deepgram Speech-to-Text v1 (/v1/listen) for prerecorded; live audio transcription; include transcribe; live transcription.

How do I install Deepgram JS Speech To Text in Claude Code?

Run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-speech-to-text -a claude-code`. Or copy the skill folder (.agents/skills/deepgram-js-speech-to-text in deepgram/deepgram-js-sdk) into .claude/skills/deepgram-js-speech-to-text in your project. Claude Code loads it when a task matches its description.

How do I install Deepgram JS Speech To Text in Codex?

Run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-speech-to-text -a codex`. Or copy the skill folder (.agents/skills/deepgram-js-speech-to-text in deepgram/deepgram-js-sdk) into .agents/skills/deepgram-js-speech-to-text in your project. Codex loads it when a task matches its description.

Can I use Deepgram JS Speech To Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepgram/deepgram-js-sdk --skill deepgram-js-speech-to-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepgram-js-speech-to-text, .gemini/skills/deepgram-js-speech-to-text, .github/skills/deepgram-js-speech-to-text and .opencode/skills/deepgram-js-speech-to-text in your project.

What does Deepgram JS Speech To Text need to run?

Going by SKILL.md and its folder, Deepgram JS Speech To Text needs the command-line tools its instructions call (npx) and credentials named DEEPGRAM_API_KEY. Our summary lists: Node.js; A credential in DEEPGRAM_API_KEY.

Does Deepgram JS Speech To Text access the network?

SKILL.md names 2 domains. In commands or code: dpgr.am; the agent is likely to contact it when it follows the instructions. As links in the text: developers.deepgram.com. This is read from the text; nothing was executed.

Is Deepgram JS Speech To Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deepgram JS Speech To Text use?

Deepgram JS Speech To Text is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepgram JS Speech To Text use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deepgram JS Speech To Text?

Skills that share tags, products or a category with Deepgram JS Speech To Text: Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars), 9Router Speech-to-Text (decolua/9router, 31k stars), Speech Engine (elevenlabs/skills, 482 stars) and Azure AI (microsoft/GitHub-Copilot-for-Azure, 255 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepgram JS Speech To Text?

deepgram (a GitHub organization) maintains it in deepgram/deepgram-js-sdk, which has 276 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: deepgram/deepgram-js-sdk on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.