Agent skill

Testing Livekit Agents

by livekit-examples in livekit-examples/agent-starter-python

Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

MITAuto-check passedTesting & QA

Install Testing Livekit Agents

skills CLI
$ npx skills add livekit-examples/agent-starter-python --skill testing-livekit-agents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install livekit-examples/agent-starter-python testing-livekit-agents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/livekit-examples/agent-starter-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/testing-livekit-agents .claude/skills/testing-livekit-agents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
testing-livekit-agents
GitHub stars
264
Used in
1 other repo
Token cost
~1.9k tokens
SKILL.md length
1,060 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

  • Works in 6 steps: The behavior the user asked for. Every… → Tool invocation: the right tool with the… → Tool failure: what the agent says when a… → …
  • The user asks to write tests for my agent
  • SKILL.md covers What a test asserts on, Mocking tools, Multi-turn and seeded history and What to test, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Testing Livekit Agents is an agent skill from livekit-examples/agent-starter-python. Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent behavior that needs regression coverage. Covers the SDK's test session harness, assertions on messages, tool calls and handoffs, LLM judging of a reply against an intent, mocking tools, multi-turn tests, and judging…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Unit testing, Test generation and Agent evaluation and testing. It works with pytest, Vitest, Python and Node.js. The repository describes itself as: A complete voice AI starter for LiveKit Agents with Python. The licence is MIT.

When your agent uses it

  • The user asks to write tests for my agent
  • Add a test for this tool
  • Test the handoff
  • Why does my agent test fail

Example prompts

  • “write tests for my agent”
  • “add a test for this tool”
  • “test the handoff”
  • “/testing-livekit-agents”

Requirements

  • Node.js

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. The behavior the user asked for. Every agent needs at least this.
  2. Tool invocation: the right tool with the right arguments for a representative request.
  3. Tool failure: what the agent says when a tool errors or returns nothing.
  4. Refusals and limits: the agent declines what it should decline and doesn't make up data it
  5. Handoffs: the transition fires when it should, and the next agent has what it needs.
  6. Every bug you've fixed. Without a test, a fixed bug can come back unnoticed.

What it can do on your machine

Read from SKILL.md and the folder at commit 76ddabb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Testing Livekit Agents loads about 1.9k tokens when it runs. Until then it costs about 178 tokens; SKILL.md has 1,060 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~178
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from livekit-examples/agent-starter-python at commit 76ddabb, republished under its MIT licence (© livekit-examples). 1,060 words, ~1,898 tokens.

Download SKILL.mdSave it as .claude/skills/testing-livekit-agents/SKILL.md (or your agent's skills folder).
name
testing-livekit-agents
description
Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent behavior that needs regression coverage. Covers the SDK's test session harness, assertions on messages, tool calls and handoffs, LLM judging of a reply against an intent, mocking tools, multi-turn tests, and judging whole conversations with the built-in judges. For interactive poking use debugging-livekit-agents. For grading whole conversations at scale use running-livekit-simulations.
license
MIT
metadata.author
livekit

Testing LiveKit agents

Turn-level tests are the cheapest lasting verification an agent can have. They run in the user's existing suite, in text mode, and they're fast enough for every commit. The framework's helper names and signatures change, so look them up with reading-livekit-docs before writing, and use the testing docs page for every API detail this skill leaves out.

Every test has the same shape: start a test session with the agent under test, run one user turn, and assert on the events that turn produced.

What a test asserts on

A turn produces a sequence of events. A simple turn is one message. A more typical one is a tool call, its output, maybe a handoff, and then a message. You write the test by walking that sequence in order, asserting on each event, and then asserting the turn has nothing more.

There are three kinds of assertion, and most of this skill is knowing which one to use.

Structural assertions check message roles, that a tool was called, its arguments, what it returned, and that a handoff to a specific agent happened. They're deterministic and fail for exactly one reason, so prefer them.

Judged assertions use an LLM-judge helper that gives one message and an intent string to a model and asks whether they match. Use them for the content of a reply, which you can't assert exactly. Describe the intent by outcome ("tells the user the booking is confirmed and gives the time"), not by wording.

Whole-conversation assertions use a judge-group helper that runs several built-in judges concurrently over the whole chat history and aggregates their verdicts. The built-in judges cover dimensions like grounding, relevance, safety, task completion, and tool use; the docs list the current set. Use this when the question spans several turns.

Rules of thumb:

  • Assert structure first and judge only what's left. A judged assertion that could have been structural is slower and flakier.
  • Close the turn. Assert there are no further events. Otherwise a test can pass while the agent also does something you didn't intend.
  • One behavior per test. When a test that asserts six things fails, it tells you very little.

Mocking tools

Tests shouldn't hit real backends, since that makes a suite slow and nondeterministic. Override the tools for the agent under test and return fixed values.

Two things beyond the API:

  • Mock failures as well as successes. Returning an error from a mock makes the tool raise, which lets you test what the agent says when a backend is down. Most agents lack this test, and it's the behavior users notice most.
  • Scope matters. By default, mocks apply in a block around your own run() calls, which is what a test needs. A session-scoped form also exists for a session that runs on its own and needs mocks active for its whole lifetime. Simulation entrypoints use that form, tests don't. See writing-livekit-scenarios.

Mocking only changes execution. The model still sees the real tool schemas, so tool selection is still under test.

Multi-turn and seeded history

There are two ways to test behavior that depends on earlier turns:

  • Run the turns. Each turn adds to the conversation history, so the second turn sees the first. Use this when the path matters: collecting details across turns, the user changing their mind mid-flow, a handoff carrying context.
  • Seed the history. Build a chat history and give it to the agent to start at the state you care about. Use this when the setup turns aren't what you're testing. It's faster and less brittle than replaying five turns to reach the sixth.
Show full SKILL.md (459 more words)Show less

What to test

Roughly in order of value:

  1. The behavior the user asked for. Every agent needs at least this.
  2. Tool invocation: the right tool with the right arguments for a representative request.
  3. Tool failure: what the agent says when a tool errors or returns nothing.
  4. Refusals and limits: the agent declines what it should decline and doesn't make up data it can't have. The refusal is the pass condition.
  5. Handoffs: the transition fires when it should, and the next agent has what it needs.
  6. Every bug you've fixed. Without a test, a fixed bug can come back unnoticed.

Three kinds of evidence, kept distinct

A test suite for an agent that does things — books, edits, confirms — needs three kinds of test, and confusing them is how a green suite ships a broken agent:

  • Deterministic application tests prove guards, no-ops, empty-vs-unknown values, corrections, storage failure, and retries, against the application core directly. Fast, exact, no model.
  • Real-SDK tests with a scripted provider prove event wiring and tool plumbing: that the listener fires, that the tool receives what the session sends. A scripted provider proves mechanics, never language understanding — its exact sentences are fixtures, not requirements.
  • Ordinary-language tests with the real model prove the model can complete the task from how callers actually talk: paraphrases absent from the prompt, corrections mixed with agreement, the same short reply answering different questions.

Shortcuts that look like tests and aren't: calling a helper or filling private state instead of driving the interaction; telling the test caller which tool to name; mocking the very mutation whose correctness is under test; comparing output only with the application's own exporter; retrying a failed turn until it passes; deleting the assertion that failed.

Don't write tests that punish correct behavior

The most common bad agent test asserts that the agent does something it shouldn't: states data it can't know, gives a specific medical, legal, or financial recommendation, or completes a flow that should have been blocked. If the agent can only pass by misbehaving, fix the test.

For a guardrail, the pass is that the agent refuses, escalates, or declines to make something up. Write the assertion that way.

Where this sits

  • Interactive poking while building → debugging-livekit-agents. Faster loop, no assertions kept.
  • Turn-level, deterministic, every commit → here. Cheapest lasting check.
  • Whole conversations graded by a simulated user, before a release → running-livekit-simulations. More expensive, and catches emergent behavior that turn-level tests miss.

When a simulation keeps failing the same way, the bug is usually turn-level. Write a test for it here, which pins down the cause more precisely and catches it earlier.

  • Current API surface: reading-livekit-docs
  • Finding the bug first: debugging-livekit-agents
  • Simulations, including seeded state and session-scoped mocks: writing-livekit-scenarios, running-livekit-simulations

© livekit-examples, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/testing-livekit-agents of livekit-examples/agent-starter-python.

Open the folder on GitHubat commit 76ddabb

Used in 2 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in livekit-examples/agent-starter-python, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Testing Livekit Agents next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Testing Livekit Agents compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Testing Livekit Agents this skilllivekit-examples/agent-starter-python2641 repos~1.9kAutomated safety check: PassMIT
TDD Guidealirezarezvani/claude-skills28k—~3.4kAutomated safety check: PassMIT
TDD GuideLeoYeAI/openclaw-master-skills2.2k—~1.4kAutomated safety check: PassMIT
TDD GuideaAAaqwq/AGI-Super-Team1052 repos~1.1kAutomated safety check: PassMIT
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence

Similar skills

  • TDD Guide

    alirezarezvani/claude-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    28k GitHub stars~3.4k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • TDD Guide

    LeoYeAI/openclaw-master-skills

    Test-driven development skill for writing unit tests, generating test fixtures and mocks, analyzing coverage gaps, and guiding red-green-refactor workflows across Jest, Pytest, JUnit, Vitest, and…

    2.2k GitHub stars~1.4k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • TDD Guide

    aAAaqwq/AGI-Super-Team

    Test-driven development workflow with test generation, coverage analysis, and multi-framework support

    105 GitHub starsUsed in 2 repos~1.1k tokens
    Testing & QAAuto-check passed
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Test Coverage Review

    areed1192/finance-news-aggregator

    Audit, plan, write, and verify unit tests for Python projects using pytest.

    149 GitHub stars~2.6k tokensUpdated 5 mo ago
    Testing & QAAuto-check passed

More from livekit-examples/agent-starter-python

  • Writing Livekit Scenarios

    livekit-examples/agent-starter-python

    Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them.

    264 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Debugging Livekit Agents

    livekit-examples/agent-starter-python

    Drives a multi-turn conversation with a LiveKit agent running locally to see what it does.

    264 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Reading Livekit Docs

    livekit-examples/agent-starter-python

    Looks up current LiveKit facts (API signatures, CLI flags, config options, model and provider support, SDK changelogs, pricing) from the docs instead of answering from memory.

    264 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Running Livekit Simulations

    livekit-examples/agent-starter-python

    Runs LiveKit agent simulations and acts on the results. An agent skill from livekit-examples/agent-starter-python.

    264 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Building Livekit Agents

    livekit-examples/agent-starter-python

    Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud.

    264 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check: notes
  • Operating Livekit Agents

    livekit-examples/agent-starter-python

    Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside…

    264 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed

Categories

Questions about Testing Livekit Agents

What does Testing Livekit Agents do?

Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js). Testing Livekit Agents is an agent skill from livekit-examples/agent-starter-python.js).

When should I use Testing Livekit Agents?

Testing Livekit Agents fits situations like: the user asks to write tests for my agent; add a test for this tool; test the handoff; why does my agent test fail.

How do I install Testing Livekit Agents in Claude Code?

Run `npx skills add livekit-examples/agent-starter-python --skill testing-livekit-agents -a claude-code`. Or copy the skill folder (.agents/skills/testing-livekit-agents in livekit-examples/agent-starter-python) into .claude/skills/testing-livekit-agents in your project. Claude Code loads it when a task matches its description.

How do I install Testing Livekit Agents in Codex?

Run `npx skills add livekit-examples/agent-starter-python --skill testing-livekit-agents -a codex`. Or copy the skill folder (.agents/skills/testing-livekit-agents in livekit-examples/agent-starter-python) into .agents/skills/testing-livekit-agents in your project. Codex loads it when a task matches its description.

Can I use Testing Livekit Agents in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add livekit-examples/agent-starter-python --skill testing-livekit-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/testing-livekit-agents, .gemini/skills/testing-livekit-agents, .github/skills/testing-livekit-agents and .opencode/skills/testing-livekit-agents in your project.

What does Testing Livekit Agents need to run?

SKILL.md names no scripts, command-line tools or credentials: Testing Livekit Agents is instructions for the agent only. Our summary lists: Node.js.

Does Testing Livekit Agents access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Testing Livekit Agents safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Testing Livekit Agents use?

Testing Livekit Agents is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Testing Livekit Agents use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Testing Livekit Agents?

Skills that share tags, products or a category with Testing Livekit Agents: TDD Guide (alirezarezvani/claude-skills, 28k stars), TDD Guide (LeoYeAI/openclaw-master-skills, 2.2k stars), TDD Guide (aAAaqwq/AGI-Super-Team, 105 stars) and Adk Verify Snippets (google/adk-python, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Testing Livekit Agents?

livekit-examples (a GitHub organization) maintains it in livekit-examples/agent-starter-python, which has 264 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.

Source: livekit-examples/agent-starter-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.