Azure AI Voicelive Py
microsoft/skills
Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).
Builds a full-duplex Python voice agent on Deepgram's agent.converse WebSocket, combining speech-to-text, an LLM and text-to-speech with interruption and function calling.
$ npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install deepgram/deepgram-python-sdk deepgram-python-voice-agent --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/deepgram/deepgram-python-sdk.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/deepgram-python-voice-agent .claude/skills/deepgram-python-voice-agent && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "deepgram-python-voice-agent" agent skill from https://github.com/deepgram/deepgram-python-sdk/tree/main/.agents/skills/deepgram-python-voice-agent into .claude/skills/deepgram-python-voice-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepgram-python-voice-agent", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/deepgram/deepgram-python-sdk/tree/main/.agents/skills/deepgram-python-voice-agentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install deepgram/deepgram-python-sdk deepgram-python-voice-agent --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/deepgram/deepgram-python-sdk.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/deepgram-python-voice-agent .agents/skills/deepgram-python-voice-agent && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "deepgram-python-voice-agent" agent skill from https://github.com/deepgram/deepgram-python-sdk/tree/main/.agents/skills/deepgram-python-voice-agent into .agents/skills/deepgram-python-voice-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepgram-python-voice-agent", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install deepgram/deepgram-python-sdk deepgram-python-voice-agent --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/deepgram/deepgram-python-sdk.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/deepgram-python-voice-agent .cursor/skills/deepgram-python-voice-agent && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "deepgram-python-voice-agent" agent skill from https://github.com/deepgram/deepgram-python-sdk/tree/main/.agents/skills/deepgram-python-voice-agent into .cursor/skills/deepgram-python-voice-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepgram-python-voice-agent", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/deepgram/deepgram-python-sdk.git --path .agents/skills/deepgram-python-voice-agent--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install deepgram/deepgram-python-sdk deepgram-python-voice-agent --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/deepgram/deepgram-python-sdk.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/deepgram-python-voice-agent .gemini/skills/deepgram-python-voice-agent && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "deepgram-python-voice-agent" agent skill from https://github.com/deepgram/deepgram-python-sdk/tree/main/.agents/skills/deepgram-python-voice-agent into .gemini/skills/deepgram-python-voice-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepgram-python-voice-agent", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install deepgram/deepgram-python-sdk deepgram-python-voice-agentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/deepgram/deepgram-python-sdk.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/deepgram-python-voice-agent .github/skills/deepgram-python-voice-agent && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "deepgram-python-voice-agent" agent skill from https://github.com/deepgram/deepgram-python-sdk/tree/main/.agents/skills/deepgram-python-voice-agent into .github/skills/deepgram-python-voice-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepgram-python-voice-agent", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install deepgram/deepgram-python-sdk deepgram-python-voice-agent --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/deepgram/deepgram-python-sdk.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/deepgram-python-voice-agent .opencode/skills/deepgram-python-voice-agent && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "deepgram-python-voice-agent" agent skill from https://github.com/deepgram/deepgram-python-sdk/tree/main/.agents/skills/deepgram-python-voice-agent into .opencode/skills/deepgram-python-voice-agent/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "deepgram-python-voice-agent", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
deepgram-python-voice-agentBuilds a full-duplex Python voice agent on Deepgram's agent.converse WebSocket, combining speech-to-text, an LLM and text-to-speech with interruption and function calling.
This skill covers Deepgram's hosted voice agent runtime, reached through `client.agent.v1.connect()` against `wss://agent.deepgram.com/v1/agent/converse` with an `Authorization: Token` header. It fits building an interactive assistant where the user can interrupt the agent mid-reply, where tool or function calls are triggered by the conversation, and where Deepgram hosts the STT, LLM and TTS orchestration instead of you wiring the three separately.
Configuration goes out first as `AgentV1Settings` through `send_settings`, audio frames follow through `send_media`, and the code reacts to server events such as `Welcome`, `SettingsApplied`, `ConversationText`, `UserStartedSpeaking`, `AgentThinking`, `FunctionCallRequest`, `AgentStartedSpeaking` and `AgentAudioDone`. Client-side messages beyond settings and media include `KeepAlive` for long sessions, mid-session prompt or voice updates, text injection and `ForceEndTurn`, the last requiring a V2/Flux listen provider. For one-way transcription, one-way synthesis or persisted agent configs, the skill points to its sibling Deepgram Python skills instead.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5c2f3af. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npxFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
developers.deepgram.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Deepgram Python Voice Agent loads about 3.6k tokens when it runs. Until then it costs about 162 tokens; SKILL.md has 753 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from deepgram/deepgram-python-sdk at commit 5c2f3af, republished under its MIT licence (© deepgram). 753 words, ~3,602 tokens.
.claude/skills/deepgram-python-voice-agent/SKILL.md (or your agent's skills folder).Full-duplex voice agent runtime: STT + LLM (think) + TTS + function calling over a single WebSocket at agent.deepgram.com/v1/agent/converse.
Use a different skill when:
deepgram-python-speech-to-text or deepgram-python-conversational-stt.deepgram-python-text-to-speech.deepgram-python-audio-intelligence.deepgram-python-management-api.from dotenv import load_dotenv
load_dotenv()
from deepgram import DeepgramClient
client = DeepgramClient()Header: Authorization: Token <api_key>. Base URL: wss://agent.deepgram.com/v1/agent/converse.
import threading, time
from deepgram.core.events import EventType
from deepgram.agent.v1.types import (
AgentV1Settings,
AgentV1SettingsAgent,
AgentV1SettingsAgentListen,
AgentV1SettingsAgentListenProvider_V1,
AgentV1SettingsAudio,
AgentV1SettingsAudioInput,
)
from deepgram.types.speak_settings_v1 import SpeakSettingsV1
from deepgram.types.speak_settings_v1provider import SpeakSettingsV1Provider_Deepgram
from deepgram.types.think_settings_v1 import ThinkSettingsV1
from deepgram.types.think_settings_v1provider import ThinkSettingsV1Provider_OpenAi
with client.agent.v1.connect() as agent:
settings = AgentV1Settings(
audio=AgentV1SettingsAudio(
input=AgentV1SettingsAudioInput(encoding="linear16", sample_rate=24000),
),
agent=AgentV1SettingsAgent(
listen=AgentV1SettingsAgentListen(
provider=AgentV1SettingsAgentListenProvider_V1(type="deepgram", model="nova-3"),
),
think=ThinkSettingsV1(
provider=ThinkSettingsV1Provider_OpenAi(
type="open_ai", model="gpt-4o-mini", temperature=0.7,
),
prompt="You are a helpful assistant. Keep replies brief.",
),
speak=SpeakSettingsV1(
provider=SpeakSettingsV1Provider_Deepgram(type="deepgram", model="aura-2-asteria-en"),
),
),
)
agent.send_settings(settings) # MUST be first message after connect
def on_message(m):
if isinstance(m, bytes):
# agent speech audio — play or append to output buffer
return
t = getattr(m, "type", "Unknown")
if t == "ConversationText":
print(f"[{getattr(m, 'role', '?')}] {getattr(m, 'content', '')}")
elif t == "UserStartedSpeaking": print(">> user speaking")
elif t == "AgentThinking": print(">> agent thinking")
elif t == "AgentStartedSpeaking": print(">> agent speaking")
elif t == "AgentAudioDone": print(">> agent done")
elif t == "FunctionCallRequest": handle_tool_call(m)
agent.on(EventType.OPEN, lambda _: print("open"))
agent.on(EventType.MESSAGE, on_message)
agent.on(EventType.CLOSE, lambda _: print("close"))
agent.on(EventType.ERROR, lambda e: print(f"err: {e}"))
def send_audio():
for chunk in mic_chunks():
agent.send_media(chunk)
threading.Thread(target=send_audio, daemon=True).start()
agent.start_listening() # blocksWelcome — connection acknowledgedSettingsApplied — your Settings acceptedConversationText — text of a turn (with role: user or assistant)UserStartedSpeaking — VAD detected userAgentThinking — LLM is workingFunctionCallRequest — tool/function call initiated by the modelAgentStartedSpeaking — TTS startingAgentAudioDone — TTS finished for this turnWarning, ErrorSettings (send first)Media (binary audio frames in declared encoding/sample_rate)KeepAlive (on long sessions)FunctionCallRequest)ForceEndTurn (end an active user turn; requires a Deepgram V2/Flux listen provider)You can persist the agent block of a Settings message server-side and reuse it by agent_id. client.voice_agent.configurations.create stores a JSON string representing the agent object only (listen / think / speak providers + prompt) — NOT the full AgentV1Settings payload. Do not send top-level Settings fields like audio to that API; those still go in the live Settings message at connect time. The returned agent_id replaces the inline agent object in future Settings messages. Managed via client.voice_agent.configurations.* — see deepgram-python-management-api.
You can change agent behavior without disconnecting by sending control messages on the live socket. Each method is available on the agent connection object (agent in the quick-start) for both sync and async clients.
from deepgram.agent.v1.types import (
AgentV1UpdatePrompt,
AgentV1UpdateSpeak,
AgentV1UpdateSpeakSpeak, # type alias accepting SpeakSettingsV1 or list
AgentV1UpdateThink,
AgentV1UpdateThinkThink, # type alias accepting ThinkSettingsV1 or list
AgentV1InjectAgentMessage,
AgentV1InjectUserMessage,
AgentV1KeepAlive,
)
from deepgram.types.speak_settings_v1 import SpeakSettingsV1
from deepgram.types.speak_settings_v1provider import SpeakSettingsV1Provider_Deepgram
from deepgram.types.think_settings_v1 import ThinkSettingsV1
from deepgram.types.think_settings_v1provider import ThinkSettingsV1Provider_OpenAi
# 1. Swap the LLM system prompt mid-conversation (e.g. escalate to a different persona)
agent.send_update_prompt(
AgentV1UpdatePrompt(prompt="You are now in expert escalation mode. Be precise and concise.")
)
# Server replies with a `PromptUpdated` event when the new prompt is in effect.
# 2. Swap the TTS voice without reconnecting (e.g. switch language or persona)
agent.send_update_speak(
AgentV1UpdateSpeak(
speak=SpeakSettingsV1(
provider=SpeakSettingsV1Provider_Deepgram(
type="deepgram", model="aura-2-luna-en",
),
),
)
)
# Server replies with a `SpeakUpdated` event.
# 3. Swap the LLM provider/model (e.g. cheaper model for follow-ups)
agent.send_update_think(
AgentV1UpdateThink(
think=ThinkSettingsV1(
provider=ThinkSettingsV1Provider_OpenAi(
type="open_ai", model="gpt-4o-mini", temperature=0.3,
),
prompt="You are a helpful assistant. Keep replies brief.",
),
)
)
# Server replies with a `ThinkUpdated` event.
# 4. Force the agent to say something specific (without waiting for user audio)
agent.send_inject_agent_message(
AgentV1InjectAgentMessage(message="Quick reminder: your call is being recorded.")
)
# Useful for proactive prompts, status updates, or scripted segues.
# 5. Inject a user message (e.g. text input from a chat sidebar alongside voice)
agent.send_inject_user_message(
AgentV1InjectUserMessage(content="Schedule a follow-up for next Tuesday at 2pm.")
)
# Server may reply with `InjectionRefused` if the agent is mid-utterance — retry after `AgentAudioDone`.
# 6. Idle-period keep-alive (no payload required; the SDK fills in the type literal)
agent.send_keep_alive(AgentV1KeepAlive())
# Or simply: agent.send_keep_alive() — the message arg is optional.
# 7. End an active user turn immediately (for example, on push-to-talk release).
# Requires a Deepgram V2/Flux listen provider; V1 returns FORCE_END_TURN_UNSUPPORTED.
agent.send_force_end_turn()Async client equivalents are identical but await-prefixed:
await agent.send_update_prompt(AgentV1UpdatePrompt(prompt="..."))
await agent.send_inject_agent_message(AgentV1InjectAgentMessage(message="..."))
await agent.send_force_end_turn()Continuous voice agents need explicit handling for idle periods, stream pauses, and reconnects.
Pause / idle (no audio for several seconds): stop calling send_media, but emit a KeepAlive every ~5 seconds. Without it, the server closes the socket at ~10 seconds of idle.
import threading, time
stop = threading.Event()
def keepalive_loop():
while not stop.is_set():
if stop.wait(5):
return
try:
agent.send_keep_alive()
except Exception:
return # socket closed; outer loop will reconnect
threading.Thread(target=keepalive_loop, daemon=True).start()Resume after pause: just call send_media again. No control message is required — the agent picks up VAD on the next chunk.
Reconnect after disconnect (preserve conversation context): Settings cannot be re-sent on the same closed socket; open a new connection and resend the same Settings. To carry conversation history forward, include it in the new Settings.agent.context.messages so the LLM resumes with prior turns:
from deepgram.agent.v1.types import (
AgentV1SettingsAgentContext,
AgentV1SettingsAgentContextMessagesItem,
AgentV1SettingsAgentContextMessagesItemContent,
AgentV1SettingsAgentContextMessagesItemContentRole,
)
# Build the new Settings with the captured prior turns
context = AgentV1SettingsAgentContext(
messages=[
AgentV1SettingsAgentContextMessagesItem(
content=AgentV1SettingsAgentContextMessagesItemContent(
role=AgentV1SettingsAgentContextMessagesItemContentRole.USER,
content="Hi, I'd like to schedule a meeting.",
),
),
AgentV1SettingsAgentContextMessagesItem(
content=AgentV1SettingsAgentContextMessagesItemContent(
role=AgentV1SettingsAgentContextMessagesItemContentRole.ASSISTANT,
content="Sure — what day works best?",
),
),
],
)
new_settings = settings.model_copy(update={"agent": settings.agent.model_copy(update={"context": context})})
# Open a fresh connection and replay
with client.agent.v1.connect() as agent2:
agent2.send_settings(new_settings)
# ... same handlers + audio loop as beforeThe server emits a History message on connect when the SDK has captured prior turns; in Python you receive this as an AgentV1History object (wire type literal: "History"). Persist these turns in your application so a reconnect can rebuild context.messages.
Detect disconnects: the EventType.CLOSE handler fires before the with block exits. Catch it and trigger your reconnect logic from there. Check EventType.ERROR payloads for cause (network drop vs server-initiated close vs warning).
reference.md — "Agent V1 Connect", "Voice Agent Configurations"./llmstxt/developers_deepgram_llms_txt.Authorization: Token <api_key>. Temporary / access tokens (created via client.auth.v1.tokens.grant() or an equivalent server) use Authorization: Bearer <access_token>. The custom DeepgramClient in this repo accepts an access_token parameter and installs a Bearer override for all HTTP + WebSocket calls — see src/deepgram/client.py.agent.deepgram.com, not api.deepgram.com.Settings IMMEDIATELY after connect — no audio before settings are applied.ThinkSettingsV1Provider_OpenAi, SpeakSettingsV1Provider_Deepgram, ...). Pick the right union variant; don't pass raw dicts.socket_client.py is temporarily frozen (see .fernignore → src/deepgram/agent/v1/socket_client.py) and currently carries _sanitize_numeric_types plus the construct_type / broad-catch fixes — needed for unknown WS message shapes. Expected to be unfrozen during a future Fern regen and re-compared.examples/30-voice-agent.pyexamples/32-voice-agent-force-end-turn.py — Force an active turn to end with a Flux listen providertests/manual/agent/v1/connect/main.py — live connection testFor cross-language Deepgram product knowledge — the consolidated API reference, documentation finder, focused runnable recipes, third-party integration examples, and MCP setup — install the central skills:
npx skills add deepgram/skillsThis SDK ships language-idiomatic code skills; deepgram/skills ships cross-language product knowledge (see api, docs, recipes, examples, starters, setup-mcp).
© deepgram, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/deepgram-python-voice-agent of deepgram/deepgram-python-sdk.
Open the folder on GitHubat commit 5c2f3af
Deepgram Python Voice Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Deepgram Python Voice Agent this skilldeepgram/deepgram-python-sdk | 469 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Azure AI Voicelive Pymicrosoft/skills | 3.1k | 6 repos | ~2.9k | Automated safety check: Pass | MIT | |
| Gemini Live API Devgoogle-gemini/gemini-skills | 4.3k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Deepgram JS Audio Intelligencedeepgram/deepgram-js-sdk | 276 | — | ~1.5k | Automated safety check: Pass | MIT | |
| Gemini Live API DevJetBrains/skills | 363 | — | ~2.6k | Automated safety check: Pass | None | |
| Speech Engineelevenlabs/skills | 479 | — | ~2.5k | Automated safety check: Warn | MIT |
microsoft/skills
Build real-time voice AI applications using Azure AI Voice Live SDK (azure-ai-voicelive).
google-gemini/gemini-skills
A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.
deepgram/deepgram-js-sdk
A skill your agent uses when writing or reviewing JavaScript/TypeScript in this repo that calls Deepgram audio analytics overlays on /v1/listen - summarize, topics, intents, sentiment, diarize…
JetBrains/skills
A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API.
elevenlabs/skills
Add real-time voice conversations to a custom agent runtime with ElevenLabs Speech Engine.
majiayu000/claude-skill-registry
Expert in building voice AI applications - from real-time voice agents to voice-enabled apps.
deepgram/deepgram-python-sdk
Shows how to add Deepgram analytics such as diarization, summaries, sentiment, topics, redaction and language detection to speech transcription in Python.
deepgram/deepgram-python-sdk
Writes and reviews Python code for Deepgram's turn-aware streaming speech-to-text (Flux, /v2/listen), including end-of-turn detection.
deepgram/deepgram-python-sdk
Guides Python code that calls the Deepgram Management APIs to administer projects, keys, members, usage, billing and stored Voice Agent configurations.
deepgram/deepgram-python-sdk
Covers basic transcription with the Deepgram Python SDK's listen.v1 endpoint, for one-shot REST transcription of a file or URL and live WebSocket streaming with interim results.
deepgram/deepgram-python-sdk
Uses the Deepgram Python SDK's Read API to analyze text for sentiment, summaries, topics and intents with client.read.v1.text.analyze, from raw text or a hosted URL.
deepgram/deepgram-python-sdk
Guides Python code that calls Deepgram Text-to-Speech v1, covering one-shot REST, streaming WebSocket and the TextBuilder helper.
Categories
Builds a full-duplex Python voice agent on Deepgram's agent.converse WebSocket, combining speech-to-text, an LLM and text-to-speech with interruption and function calling. com/v1/agent/converse` with an `Authorization: Token` header. It fits building an interactive assistant where the user can interrupt the agent mid-reply, where tool or function calls are triggered by the conversation, and where Deepgram hosts the STT, LLM and TTS orchestration instead of you wiring the three separately.
Deepgram Python Voice Agent fits situations like: building a voice assistant that users can interrupt while it speaks; adding function or tool calling triggered by a spoken conversation; wiring the AgentV1Settings and event handlers for a Deepgram voice agent; choosing a Deepgram skill for text-to-speech or transcription instead of a full agent.
Run `npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a claude-code`. Or copy the skill folder (.agents/skills/deepgram-python-voice-agent in deepgram/deepgram-python-sdk) into .claude/skills/deepgram-python-voice-agent in your project. Claude Code loads it when a task matches its description.
Run `npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a codex`. Or copy the skill folder (.agents/skills/deepgram-python-voice-agent in deepgram/deepgram-python-sdk) into .agents/skills/deepgram-python-voice-agent in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add deepgram/deepgram-python-sdk --skill deepgram-python-voice-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepgram-python-voice-agent, .gemini/skills/deepgram-python-voice-agent, .github/skills/deepgram-python-voice-agent and .opencode/skills/deepgram-python-voice-agent in your project.
Going by SKILL.md and its folder, Deepgram Python Voice Agent needs the command-line tools its instructions call (npx). Our summary lists: Python with the Deepgram SDK; A Deepgram API key.
SKILL.md names 1 domain. As links in the text: developers.deepgram.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Deepgram Python Voice Agent is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Deepgram Python Voice Agent: Azure AI Voicelive Py (microsoft/skills, 3.1k stars), Gemini Live API Dev (google-gemini/gemini-skills, 4.3k stars), Deepgram JS Audio Intelligence (deepgram/deepgram-js-sdk, 276 stars) and Gemini Live API Dev (JetBrains/skills, 363 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
deepgram (a GitHub organization) maintains it in deepgram/deepgram-python-sdk, which has 469 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.
Source: deepgram/deepgram-python-sdk on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.