Agent skill

Deepgram Flux Conversational STT

by deepgram in deepgram/deepgram-python-sdk

Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection.

MITAuto-check passedAI & LLM Engineering

Install Deepgram Flux Conversational STT

skills CLI
$ npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-conversational-stt -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install deepgram/deepgram-python-sdk deepgram-python-conversational-stt --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/deepgram/deepgram-python-sdk.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deepgram-python-conversational-stt .claude/skills/deepgram-python-conversational-stt && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepgram-python-conversational-stt
GitHub stars
469
Token cost
~1.8k tokens
SKILL.md length
476 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection.

  • Works in 4 steps: In-repo reference: reference.md —… → AsyncAPI (WSS):… → Context7: library ID… → …
  • Building a conversational UI that needs explicit turn boundaries
  • SKILL.md covers When to use this product, Authentication, Quick start and Key parameters, plus 7 more sections
  • Calls npx; needs DEEPGRAM_API_KEY

What it does

The skill covers the `client.listen.v2.connect` call, a websocket-only endpoint with no REST path, built for conversational audio with explicit turn boundaries, eager end-of-turn signals and barge-in scenarios. It requires a Flux model, `flux-general-en` for English or `flux-general-multi` for multilingual, and v2 has no `language` parameter, only a `language_hint` for the multilingual model.

Key parameters include `encoding`, `sample_rate` (passed as a string), `eager_eot_threshold`, `eot_threshold`, `eot_timeout_ms` and `keyterm`. For application-controlled turns, set `eot_threshold` to the string `1.0` with a large timeout and call `send_force_end_turn()`, which needs deployment enablement. The skill sends general transcription to `deepgram-python-speech-to-text`, interactive voice agents to `deepgram-python-voice-agent` and analytics to `deepgram-python-audio-intelligence`, and its authentication example loads the key from the environment.

When your agent uses it

  • Building a conversational UI that needs explicit turn boundaries
  • Streaming transcription with Flux models and low-latency end-of-turn signals
  • Reviewing Python code that calls listen.v2

Example prompts

  • “Stream microphone audio to Deepgram Flux and print each completed turn.”
  • “Tune eager end-of-turn so my voice bot responds sooner, then handle barge-in.”
  • “Review this listen.v2 connection code, I think the model name is wrong.”

Requirements

  • A Deepgram API key
  • Python with the Deepgram SDK installed

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. In-repo reference: reference.md — "Listen V2 Connect".
  2. AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  3. Context7: library ID /llmstxt/developers_deepgram_llms_txt.
  4. Product docs

What it can do on your machine

Read from SKILL.md and the folder at commit 5c2f3af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • developers.deepgram.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • DEEPGRAM_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Deepgram Flux Conversational STT loads about 1.8k tokens when it runs. Until then it costs about 138 tokens; SKILL.md has 476 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~138
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from deepgram/deepgram-python-sdk at commit 5c2f3af, republished under its MIT licence (© deepgram). 476 words, ~1,792 tokens.

Download SKILL.mdSave it as .claude/skills/deepgram-python-conversational-stt/SKILL.md (or your agent's skills folder).
name
deepgram-python-conversational-stt
description
Use when writing or reviewing Python code in this repo that calls Deepgram Conversational STT v2 / Flux (`/v2/listen`) for turn-aware streaming transcription. Covers `client.listen.v2.connect(...)`, Flux models, end-of-turn detection. Use `deepgram-python-speech-to-text` for standard v1 ASR, `deepgram-python-voice-agent` for full-duplex interactive assistants. Triggers include "flux", "v2 listen", "conversational STT", "turn detection", "end of turn", "EOT", "listen.v2", "flux-general-en", "flux-general-multi".

Using Deepgram Conversational STT / Flux (Python SDK)

Turn-aware streaming STT at /v2/listen — optimized for conversational audio (end-of-turn detection, eager EOT, barge-in scenarios).

When to use this product

  • You're building a conversational UI and need explicit turn boundaries.
  • You want Flux models (optimized for human-to-human or human-to-agent conversation).
  • You want lower latency turn signals than v1 utterance_end.

Use a different skill when:

  • You want general-purpose transcription (captions, batch, non-conversational) → deepgram-python-speech-to-text.
  • You want a full interactive agent (STT + LLM + TTS) → deepgram-python-voice-agent.
  • You want analytics (summarize/sentiment) → deepgram-python-audio-intelligence.

Authentication

python
import os
from dotenv import load_dotenv
load_dotenv()

from deepgram import DeepgramClient
client = DeepgramClient(api_key=os.environ["DEEPGRAM_API_KEY"])

Header: Authorization: Token <api_key>. WSS only — no REST path on v2.

Quick start

python
import threading, time
from pathlib import Path
from deepgram.core.events import EventType
from deepgram.listen.v2.types import (
    ListenV2CloseStream,
    ListenV2Connected,
    ListenV2FatalError,
    ListenV2TurnInfo,
)

with client.listen.v2.connect(
    model="flux-general-en",
    encoding="linear16",
    sample_rate="16000",
) as conn:

    def on_message(m):
        if isinstance(m, ListenV2TurnInfo):
            print(f"turn {m.turn_index} [{m.event}] {m.transcript}")
        elif isinstance(m, dict):                     # untyped fallback
            if m.get("type") == "TurnInfo":
                print(f"turn {m.get('turn_index')} [{m.get('event')}] {m.get('transcript')}")
        else:
            print(f"event: {getattr(m, 'type', type(m).__name__)}")

    conn.on(EventType.OPEN,    lambda _: print("open"))
    conn.on(EventType.MESSAGE, on_message)
    conn.on(EventType.CLOSE,   lambda _: print("close"))
    conn.on(EventType.ERROR,   lambda e: print(f"err: {type(e).__name__}: {e}"))

    def send_audio():
        for chunk in mic_chunks():                     # 80ms recommended
            conn.send_media(chunk)
            time.sleep(0.01)
        conn.send_close_stream(ListenV2CloseStream(type="CloseStream"))

    threading.Thread(target=send_audio, daemon=True).start()
    conn.start_listening()

Key parameters

ParamNotes
modelflux-general-en (English) or flux-general-multi (multilingual) — REQUIRED, must be a Flux model
encodinglinear16, mulaw, etc. Omit for containerized audio
sample_rateString in the SDK signature, e.g. "16000"
eager_eot_thresholdFire end-of-turn early at this confidence
eot_thresholdPrimary end-of-turn confidence; set to "1.0" to suppress confidence-based endings
eot_timeout_msTime-based fallback turn end; still applies when eot_threshold="1.0"
keytermBias for domain keywords
mip_opt_out, tagMetadata / privacy flags
language_hintONLY for flux-general-multi
authorization, request_optionsOverride auth or request options

No language parameter on v2 — language is implied by model (flux-general-en) or hinted via language_hint on multi.

For application-controlled turns, use eot_threshold="1.0" with a sufficiently large eot_timeout_ms, then call conn.send_force_end_turn() for the active turn. ForceEndTurn requires deployment enablement; see examples/16-transcription-force-end-turn.py.

Events (server → client)

  • ListenV2Connected — connection established
  • ListenV2ConfigureSuccess / ListenV2ConfigureFailure — mid-session config changes
  • ListenV2TurnInfo — per-turn transcript + event (Update, EndOfTurn, EagerEndOfTurn, ...) + turn_index
  • ListenV2FatalError — terminal error

Client messages: ListenV2Media, ListenV2Configure, ListenV2ForceEndTurn, ListenV2CloseStream.

Async equivalent

python
from deepgram import AsyncDeepgramClient
client = AsyncDeepgramClient()

async with client.listen.v2.connect(model="flux-general-en", ...) as conn:
    # same .on(...) handlers, then:
    await conn.start_listening()

API reference (layered)

  1. In-repo reference: reference.md — "Listen V2 Connect".
  2. AsyncAPI (WSS): https://developers.deepgram.com/asyncapi.yaml
  3. Context7: library ID /llmstxt/developers_deepgram_llms_txt.
  4. Product docs:
Show full SKILL.md (203 more words)Show less

Gotchas

  1. /v2/listen, not /v1/listen. Different route, different client path (listen.v2 vs listen.v1).
  2. Flux models only. nova-3, base, etc. will be rejected. Use flux-general-en or flux-general-multi.
  3. No language parameter. Language is set by model choice. Use language_hint on flux-general-multi.
  4. sample_rate is a STRING in the SDK (e.g. "16000").
  5. Send ~80ms audio chunks for best turn-detection latency.
  6. Close with send_close_stream(ListenV2CloseStream(type="CloseStream")) — not send_finalize (that's v1).
  7. Messages may arrive as typed objects OR raw dicts — the SDK uses a tagged union with construct_type for unknowns. Handle both branches (see socket_client.py patch in .fernignore).
  8. socket_client.py is patched / frozen (see .fernignore → src/deepgram/listen/v2/socket_client.py). Don't overwrite that manual patch during regeneration; treat other listen/v2 files as generated unless the regen workflow says otherwise.
  9. Omit encoding/sample_rate for containerized audio (WAV, OGG, etc.) — the server detects them from the container.

Example files in this repo

  • examples/14-transcription-live-websocket-v2.py
  • tests/manual/listen/v2/connect/main.py
  • deepgram-python-speech-to-text — v1 general-purpose STT (REST + WSS)
  • deepgram-python-voice-agent — full interactive assistant

Central product skills

For cross-language Deepgram product knowledge — the consolidated API reference, documentation finder, focused runnable recipes, third-party integration examples, and MCP setup — install the central skills:

bash
npx skills add deepgram/skills

This SDK ships language-idiomatic code skills; deepgram/skills ships cross-language product knowledge (see api, docs, recipes, examples, starters, setup-mcp).

© deepgram, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/deepgram-python-conversational-stt of deepgram/deepgram-python-sdk.

Open the folder on GitHubat commit 5c2f3af

Compare with similar skills

Deepgram Flux Conversational STT next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepgram Flux Conversational STT compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepgram Flux Conversational STT this skilldeepgram/deepgram-python-sdk469—~1.8kAutomated safety check: PassMIT
Deepgram JS Conversational Sttdeepgram/deepgram-js-sdk276—~1.3kAutomated safety check: PassMIT
Deepgram JS Voice Agentdeepgram/deepgram-js-sdk276—~1.6kAutomated safety check: PassMIT
Audio To Textgodot-fun/gai181—~707Automated safety check: PassMIT
Transcribe Anythingswyxio/skills172—~8.5kAutomated safety check: PassMIT
Whisper Speech RecognitionOrchestra-Research/AI-Research-SKILLs13k8 repos~1.9kAutomated safety check: NotesMIT

Similar skills

  • Deepgram JS Conversational Stt

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram Conversational STT v2 / Flux (/v2/listen) for turn-aware streaming transcription.

    276 GitHub stars~1.3k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Deepgram JS Voice Agent

    deepgram/deepgram-js-sdk

    A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that builds an interactive voice agent via agent.deepgram.com/v1/agent/converse.

    276 GitHub stars~1.6k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Audio To Text

    godot-fun/gai

    Transcribes a single PCM16 WAV file to plain UTF-8 text locally through a standard-library Python wrapper around native SenseVoice Small F16 GGUF FunASR llama.cpp runtimes, preferring cross-vendor…

    181 GitHub stars~707 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Transcribe Anything

    swyxio/skills

    Transcribes audio and video files to text using pluggable ASR backends.

    172 GitHub stars~8.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 8 repos~1.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    30k GitHub stars~745 tokensUpdated 6 days ago
    Media & CreativeAuto-check passed

More from deepgram/deepgram-python-sdk

  • Deepgram Audio Intelligence for Python

    deepgram/deepgram-python-sdk

    Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python.

    469 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Deepgram Python Management API

    deepgram/deepgram-python-sdk

    Guides Python code that calls the Deepgram Management APIs to administer projects, keys, members, usage, billing and stored Voice Agent configurations.

    469 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Deepgram Python Speech-to-Text

    deepgram/deepgram-python-sdk

    Covers basic transcription with the Deepgram Python SDK's listen.v1 endpoint, for one-shot REST transcription of a file or URL and live WebSocket streaming with interim results.

    469 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Deepgram Text Intelligence

    deepgram/deepgram-python-sdk

    Uses the Deepgram Python SDK's Read API to analyze text for sentiment, summaries, topics and intents with client.read.v1.text.analyze, from raw text or a hosted URL.

    469 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • Deepgram Python Text-to-Speech

    deepgram/deepgram-python-sdk

    Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper.

    469 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Deepgram Python Voice Agent

    deepgram/deepgram-python-sdk

    Builds a full-duplex Python voice agent on Deepgram's agent.converse WebSocket, combining speech-to-text, an LLM and text-to-speech with interruption and function calling.

    469 GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Deepgram Flux Conversational STT

What does Deepgram Flux Conversational STT do?

Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection. connect` call, a websocket-only endpoint with no REST path, built for conversational audio with explicit turn boundaries, eager end-of-turn signals and barge-in scenarios. It requires a Flux model, `flux-general-en` for English or `flux-general-multi` for multilingual, and v2 has no `language` parameter, only a `language_hint` for the multilingual model.

When should I use Deepgram Flux Conversational STT?

Deepgram Flux Conversational STT fits situations like: building a conversational UI that needs explicit turn boundaries; streaming transcription with Flux models and low-latency end-of-turn signals; reviewing Python code that calls listen.v2.

How do I install Deepgram Flux Conversational STT in Claude Code?

Run `npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-conversational-stt -a claude-code`. Or copy the skill folder (.agents/skills/deepgram-python-conversational-stt in deepgram/deepgram-python-sdk) into .claude/skills/deepgram-python-conversational-stt in your project. Claude Code loads it when a task matches its description.

How do I install Deepgram Flux Conversational STT in Codex?

Run `npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-conversational-stt -a codex`. Or copy the skill folder (.agents/skills/deepgram-python-conversational-stt in deepgram/deepgram-python-sdk) into .agents/skills/deepgram-python-conversational-stt in your project. Codex loads it when a task matches its description.

Can I use Deepgram Flux Conversational STT in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-conversational-stt -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepgram-python-conversational-stt, .gemini/skills/deepgram-python-conversational-stt, .github/skills/deepgram-python-conversational-stt and .opencode/skills/deepgram-python-conversational-stt in your project.

What does Deepgram Flux Conversational STT need to run?

Going by SKILL.md and its folder, Deepgram Flux Conversational STT needs the command-line tools its instructions call (npx) and credentials named DEEPGRAM_API_KEY. Our summary lists: A Deepgram API key; Python with the Deepgram SDK installed.

Does Deepgram Flux Conversational STT access the network?

SKILL.md names 1 domain. As links in the text: developers.deepgram.com. This is read from the text; nothing was executed.

Is Deepgram Flux Conversational STT safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Deepgram Flux Conversational STT use?

Deepgram Flux Conversational STT is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepgram Flux Conversational STT use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Deepgram Flux Conversational STT?

Skills that share tags, products or a category with Deepgram Flux Conversational STT: Deepgram JS Conversational Stt (deepgram/deepgram-js-sdk, 276 stars), Deepgram JS Voice Agent (deepgram/deepgram-js-sdk, 276 stars), Audio To Text (godot-fun/gai, 181 stars) and Transcribe Anything (swyxio/skills, 172 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepgram Flux Conversational STT?

deepgram (a GitHub organization) maintains it in deepgram/deepgram-python-sdk, which has 469 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.

Source: deepgram/deepgram-python-sdk on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.