Agent skill

Gradbot Voice Agent

by gradium-ai in gradium-ai/gradbot

Build complete, working voice agent applications using the gradbot framework.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Gradbot Voice Agent

skills CLI
$ npx skills add gradium-ai/gradbot --skill gradbot-voice-agent -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gradium-ai/gradbot gradbot-voice-agent --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gradium-ai/gradbot.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gradbot-voice-agent .claude/skills/gradbot-voice-agent && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gradbot-voice-agent
GitHub stars
127
Token cost
~2.9k tokens
SKILL.md length
1,295 words
Files
3 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
Apache-2.0

At a glance

Build complete, working voice agent applications using the gradbot framework.

  • Works in 7 steps: Understand the Use Case → Create the App Directory → Write the Backend (main.py) → …
  • User asks to build a voice assistant
  • SKILL.md covers Important, Two Complexity Levels, Workflow and Common Patterns, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Gradbot Voice Agent is an agent skill from gradium-ai/gradbot. Build complete, working voice agent applications using the gradbot framework. Use when user asks to "build a voice assistant", "create a voice agent", "make a voice app", "build a haggling game", or any voice-interactive application with speech. Generates a FastAPI backend with STT/LLM/TTS orchestration, tool calling, system prompts, and a polished frontend UI. Works for any domain - customer service, games, booking, tutoring, roleplay, etc. Also use when user is working with an existing gradbot demo and wants to…

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/backend-template.md` and `references/frontend-guidelines.md`).

It sits in AI & LLM Engineering, covering Speech recognition and synthesis and Prompt engineering. It works with FastAPI. The repository describes itself as: Open source framework to vibecode and prototype voice agents with Gradium APIs. The licence is Apache-2.0.

When your agent uses it

  • User asks to build a voice assistant
  • Create a voice agent
  • Make a voice app
  • Build a haggling game

Example prompts

  • “build a voice assistant”
  • “create a voice agent”
  • “make a voice app”
  • “/gradbot-voice-agent”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Understand the Use Case
  2. Create the App Directory
  3. Write the Backend (main.py)
  4. Write System Prompts (tool-using agents only)
  5. Build the Frontend
  6. Create pyproject.toml
  7. Verify

What it can do on your machine

Read from SKILL.md and the folder at commit db64d77. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are toml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gradbot Voice Agent loads about 2.9k tokens when it runs, and up to ~8.5k if it reads all its reference files. Until then it costs about 143 tokens; SKILL.md has 1,295 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~143
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gradium-ai/gradbot at commit db64d77, republished under its Apache-2.0 licence (© gradium-ai). 1,295 words, ~2,890 tokens.

Download SKILL.mdSave it as .claude/skills/gradbot-voice-agent/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
gradbot-voice-agent
description
Build complete, working voice agent applications using the gradbot framework. Use when user asks to "build a voice assistant", "create a voice agent", "make a voice app", "build a haggling game", or any voice-interactive application with speech. Generates a FastAPI backend with STT/LLM/TTS orchestration, tool calling, system prompts, and a polished frontend UI. Works for any domain - customer service, games, booking, tutoring, roleplay, etc. Also use when user is working with an existing gradbot demo and wants to modify, extend, or understand it.

Gradbot Voice Agent Builder

Build working voice agent apps using the gradbot framework. Output is a complete, runnable app with backend, frontend, prompts, and configuration.

Important

  • ALWAYS read references/backend-template.md and references/frontend-guidelines.md before generating code
  • The generated app MUST follow the exact patterns from the reference files — these are battle-tested
  • Use the frontend-design skill to generate the static/index.html — do NOT write a generic HTML file

Two Complexity Levels

Simple chat (no tools)

For apps that are purely conversational with no actions/side-effects:

  • No tools, no tool handler, no state dataclass needed
  • Do NOT pass on_tool_call to gradbot.websocket.handle_session() (omit it entirely)
  • Prompt can come from a file OR from the frontend (editable textarea)
  • No prompts/ directory needed if prompt comes from frontend
  • Pass with_voices=True to gradbot.routes.setup() to enable voice selection UI
  • See the "Minimal (no tools)" template in references/backend-template.md
Tool-using agent (with tools)

For apps with actions (ordering, searching, game mechanics, etc.):

  • Define tools via gradbot.ToolDef and implement tool handlers
  • Track state in a dataclass, update it in tool handlers
  • System prompts loaded from prompts/ files
  • Include on_tool_call callback in gradbot.websocket.handle_session()
  • See the "With tools" template in references/backend-template.md

Workflow

Step 1: Understand the Use Case

Before writing code, identify:

  1. Domain: What is the voice agent about?
  2. Persona: Who is the AI character? What personality?
  3. Tools needed: What actions can it perform? (If none, use simple chat pattern)
  4. State: What data to track across the conversation? (If none, use simple chat)
  5. UI needs: What should the frontend show besides transcript?
Step 2: Create the App Directory

Simple chat (no tools, prompt from frontend):

<app-dir>/
  main.py
  static/
    index.html
  pyproject.toml

Tool-using agent (tools, file-based prompts):

<app-dir>/
  main.py
  game.py          # domain module (state, tools, tool handler)
  prompts/
    main.txt       # (or base.txt + phase1.txt, phase2.txt, etc.)
  static/
    index.html
  pyproject.toml

If the app has static data (menus, inventories, etc.), include it as JSON files.

Step 3: Write the Backend (main.py)

Consult references/backend-template.md for exact code patterns. Choose the right template:

  • Minimal (no tools): ~30 lines in a single main.py, voice selection, frontend-editable prompt
  • With tools: thin main.py (~20 lines) + domain module (game.py, ~100+ lines) with state, tools, prompts, and tool handler

Key setup rules:

  • Use gradbot.config.load(Path(__file__).parent) or gradbot.config.from_env() for config loading — it loads config.yaml from the app directory and any shared parent config automatically. Returns a Config object with .client_kwargs and .session_kwargs properties.
  • ALWAYS call gradbot.routes.setup(app, config=cfg, static_dir=...) to serve frontend and bundled JS
  • Pass with_voices=True to gradbot.routes.setup() if the frontend needs voice selection (registers /api/voices)
  • rewrite_rules enables language-specific text rewriting before TTS. Get it from voice.language.rewrite_rules (returns "en", "fr", etc.). Do NOT use .value — Lang is not a Python enum.
  • ALWAYS set silence_timeout_s to 0.0 — the default 5s causes the agent to re-prompt itself with its last message when the user is silent
  • Pass config=cfg to handle_session() to auto-set run_kwargs, output_format, and debug from the config. Or pass them individually for custom setups.

Critical rules for tool definitions (when using tools):

  • parameters_json must be a JSON string — use json.dumps()
  • NEVER use "type": "array" in parameters (some LLMs like Gemma fail). Use "type": "string" with "description": "Comma-separated list" instead
  • Tool descriptions should say WHEN to call the tool, not just what it does

Critical rules for tool handlers (when using tools):

  • on_tool_call receives 3 args: (handle, input_handle, websocket) where handle is a gradbot.ToolHandle
  • handle.name gives the tool name, handle.args gives parsed args (already deserialized dict)
  • Send results via handle.send_json({...}) (auto-serializes) or handle.send(json.dumps({...})) for raw JSON
  • Send errors via handle.send_error("message")
  • Send UI updates via websocket.send_json({"type": "custom_event", ...})
  • Use input_handle.send_config(new_config) to swap prompts/tools mid-session

Critical rules for system prompts:

  • Keep responses SHORT: "1-2 sentences max" for voice (long text = slow TTS)
  • If using tools: include "Call tools silently FIRST, then speak"
  • Be explicit about what the agent should NOT do
Step 4: Write System Prompts (tool-using agents only)

Skip this step for simple chat apps where the prompt comes from the frontend.

For tool-using agents, create prompts/base.txt and prompts/main.txt:

  • base.txt: personality, speaking style, boundaries, response length
  • main.txt: conversation flow, when to call tools, error handling

For multi-phase apps, create phase1.txt, phase2.txt, etc. and swap via input_handle.send_config().

Step 5: Build the Frontend

IMPORTANT: Invoke the frontend-design skill to create static/index.html.

When invoking frontend-design, provide these requirements:

  • It's a voice agent UI — primary interaction is speech, not typing
  • Must include a mic/call button to start/stop the session
  • Must show a live transcript (user bubbles + agent bubbles)
  • CRITICAL: Must load audio via three script tags (NOT ES module imports): opus-encoder.js, audio-processor.js, synced-audio-player.js from /static/js/ — then use SyncedAudioPlayer as a global (it is NOT an ES module)
  • Must follow the WebSocket protocol in references/frontend-guidelines.md
  • Must include echo cancellation checkbox (checked by default) — without it the agent hears its own TTS
  • For simple chat: include voice selector grid, editable prompt textarea, speed slider
  • For tool-using agents: include domain-specific content panels (menu, inventory, results, etc.)
Show full SKILL.md (520 more words)Show less
Step 6: Create pyproject.toml
toml
[project]
name = "<app-name>"
version = "0.1.0"
description = "<description>"
requires-python = ">=3.12"
dependencies = ["gradbot"]

gradbot includes fastapi, uvicorn with websocket support, pydantic-settings, and pyyaml. Add extra dependencies only if needed.

Step 7: Verify

After generating all files:

  • main.py imports gradbot (uses gradbot.websocket, gradbot.routes, gradbot.config)
  • Uses gradbot.config.load(Path(__file__).parent) or gradbot.config.from_env() for configuration
  • gradbot.routes.setup() called with static_dir and with_voices=True (if voice selection needed)
  • rewrite_rules uses voice.language.rewrite_rules (not .value)
  • silence_timeout_s set to 0.0
  • WebSocket endpoint calls gradbot.websocket.handle_session() with correct callbacks
  • If no tools: on_tool_call is NOT passed to handle_session()
  • If tools: parameters_json is a JSON string via json.dumps(), tool results via handle.send_json() or handle.send(json.dumps())
  • If tools: on_tool_call takes 3 args (handle, input_handle, websocket), uses handle.name and handle.args
  • Frontend loads JS via three <script> tags, NOT ES module imports
  • Frontend onText destructures a single object: ({ text, turnIdx, isUser }) =>
  • Frontend has echo cancellation checkbox wired to SyncedAudioPlayer
  • System prompt enforces short responses for voice

Common Patterns

Simple Voice Chat (no tools)
  • No state, no tools, no prompts/ directory
  • Voice selection + editable prompt from frontend
  • Frontend: transcript + voice grid + speed slider
  • ~40 lines of backend code
Voice Game (haggling, trivia, roleplay)
  • State tracks game progress, scores, inventory
  • Tools: game actions (buy, sell, attack, answer)
  • Prompt: character personality + game rules + win/lose conditions
  • Frontend: game state display (inventory, score, health bar)
Customer Service Agent (ordering, booking, support)
  • State tracks order/booking details
  • Tools: CRUD operations (add, remove, modify, confirm)
  • Prompt: service persona + menu/catalog knowledge + ordering rules
  • Frontend: order summary, menu display, confirmation
Tutoring / Language Learning
  • State tracks lesson progress, mistakes, topics covered
  • Tools: check answer, advance lesson, provide hint
  • Prompt: teacher persona + curriculum + encouragement style
  • Frontend: lesson content, progress tracker, exercise display
  • State tracks search results and selections
  • Tools: search, get details, book/purchase
  • Prompt: phases that swap as user progresses through workflow
  • Frontend: search results cards, detail view, confirmation

Troubleshooting

Agent gives long responses

Add to system prompt: "CRITICAL: Keep ALL responses to 1-2 SHORT sentences. This is a voice conversation, not a text chat."

Agent repeats itself / re-prompts when user is silent

Set silence_timeout_s = 0.0 in the session config. The default 5s timeout causes the agent to re-send its last message as context and generate a new response.

Tool calls fail with "Invalid JSON"

tool_handle.send() requires a valid JSON string. Always use json.dumps({...}), never pass a raw string.

Tool calls fail silently

Ensure on_tool_call is passed to handle_session(). Check that handle.send() is called with a JSON string, or use handle.send_json() with a dict.

Audio doesn't play

Verify frontend loads JS via three script tags (opus-encoder.js, audio-processor.js, synced-audio-player.js), NOT via ES module import. SyncedAudioPlayer is a global.

Transcript shows [object Object]

The onText callback receives a single object, not separate args. Must destructure: ({ text, turnIdx, isUser }) =>.

Agent hears its own voice (feedback loop)

Add an echo cancellation checkbox and wire it to SyncedAudioPlayer({ echoCancellation: checkbox.checked }).

Voice selection not working

Ensure gradbot.routes.setup(app, ..., with_voices=True) is passed. Without with_voices=True, the /api/voices endpoint is not registered.

WebSocket returns 404 / "Unsupported upgrade request"

Uvicorn needs the websockets library. Ensure you're using gradbot[demos] (which includes uvicorn[standard]). If still failing, add websockets explicitly to your dependencies.

Wrong language TTS pronunciation

Set rewrite_rules to voice.language.rewrite_rules. This enables language-specific text rewriting before synthesis.

© gradium-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/gradbot-voice-agent of gradium-ai/gradbot.

  • SKILL.md
  • references/backend-template.md
  • references/frontend-guidelines.md

Open the folder on GitHubat commit db64d77

Compare with similar skills

Gradbot Voice Agent next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gradbot Voice Agent compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gradbot Voice Agent this skillgradium-ai/gradbot127—~2.9kAutomated safety check: PassApache-2.0
Lilly Community Researchssaaffaakk/Lilly171—~1.4kAutomated safety check: PassMIT
Deps Bumplkmeta/txtify135—~585Automated safety check: PassApache-2.0
Verifylkmeta/txtify135—~583Automated safety check: PassApache-2.0
Prompt Improverseverity1/claude-code-prompt-improver1.9k1 repos~1.7kAutomated safety check: PassMIT
TriageTalAter/annyang6.8k—~810Automated safety check: NotesMIT

Similar skills

  • Lilly Community Research

    ssaaffaakk/Lilly

    Lilly community-research skill. An agent skill from ssaaffaakk/Lilly.

    171 GitHub stars~1.4k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Deps Bump

    lkmeta/txtify

    Safely update Txtify dependencies or resolve Dependabot alerts.

    135 GitHub stars~585 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Verify

    lkmeta/txtify

    Verify a Txtify change end-to-end. An agent skill from lkmeta/txtify.

    135 GitHub stars~583 tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Prompt Improver

    severity1/claude-code-prompt-improver

    This skill enriches vague prompts with targeted research and clarification before execution.

    1.9k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Triage

    TalAter/annyang

    Triage and close GitHub issues on TalAter/annyang. An agent skill from TalAter/annyang.

    6.8k GitHub stars~810 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check: notes
  • Yichen Asr

    mcncarl/yichen-skills

    逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…

    4.4k GitHub stars~780 tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Gradbot Voice Agent

What does Gradbot Voice Agent do?

Build complete, working voice agent applications using the gradbot framework. Gradbot Voice Agent is an agent skill from gradium-ai/gradbot. Build complete, working voice agent applications using the gradbot framework.

When should I use Gradbot Voice Agent?

Gradbot Voice Agent fits situations like: user asks to build a voice assistant; create a voice agent; make a voice app; build a haggling game.

How do I install Gradbot Voice Agent in Claude Code?

Run `npx skills add gradium-ai/gradbot --skill gradbot-voice-agent -a claude-code`. Or copy the skill folder (skills/gradbot-voice-agent in gradium-ai/gradbot) into .claude/skills/gradbot-voice-agent in your project. Claude Code loads it when a task matches its description.

How do I install Gradbot Voice Agent in Codex?

Run `npx skills add gradium-ai/gradbot --skill gradbot-voice-agent -a codex`. Or copy the skill folder (skills/gradbot-voice-agent in gradium-ai/gradbot) into .agents/skills/gradbot-voice-agent in your project. Codex loads it when a task matches its description.

Can I use Gradbot Voice Agent in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gradium-ai/gradbot --skill gradbot-voice-agent -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gradbot-voice-agent, .gemini/skills/gradbot-voice-agent, .github/skills/gradbot-voice-agent and .opencode/skills/gradbot-voice-agent in your project.

What does Gradbot Voice Agent need to run?

SKILL.md names no scripts, command-line tools or credentials: Gradbot Voice Agent is instructions for the agent only. Our summary lists: Python 3.

Does Gradbot Voice Agent access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Gradbot Voice Agent safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gradbot Voice Agent use?

Gradbot Voice Agent is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gradbot Voice Agent use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.6k tokens, read only when the agent opens those files.

What are the alternatives to Gradbot Voice Agent?

Skills that share tags, products or a category with Gradbot Voice Agent: Lilly Community Research (ssaaffaakk/Lilly, 171 stars), Deps Bump (lkmeta/txtify, 135 stars), Verify (lkmeta/txtify, 135 stars) and Prompt Improver (severity1/claude-code-prompt-improver, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gradbot Voice Agent?

gradium-ai (a GitHub organization) maintains it in gradium-ai/gradbot, which has 127 GitHub stars. The repository was last updated on September 19, 2026.

Source: gradium-ai/gradbot on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.