Agent skill

Running Livekit Simulations

by livekit-examples in livekit-examples/agent-starter-python

Runs LiveKit agent simulations and acts on the results. An agent skill from livekit-examples/agent-starter-python.

MITAuto-check passedTesting & QA

Install Running Livekit Simulations

skills CLI
$ npx skills add livekit-examples/agent-starter-python --skill running-livekit-simulations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install livekit-examples/agent-starter-python running-livekit-simulations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/livekit-examples/agent-starter-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/running-livekit-simulations .claude/skills/running-livekit-simulations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
running-livekit-simulations
GitHub stars
264
Used in
1 other repo
Token cost
~1.7k tokens
SKILL.md length
880 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Runs LiveKit agent simulations and acts on the results. An agent skill from livekit-examples/agent-starter-python.

  • Works in 3 steps: A bug in the agent. Fix the agent and… → A bad scenario. The expectation requires… → A gap in the simulated world. The…
  • The user says run my simulations
  • SKILL.md covers When to reach for a simulation, Running, Text or audio and Automating a pre-release run, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Running Livekit Simulations is an agent skill from livekit-examples/agent-starter-python. Runs LiveKit agent simulations and acts on the results. Use when the user says "run my simulations", "regression test my agent before deploying", "run the scenarios", "use lk agent simulate", "did my agent pass", "why did this scenario fail", "run simulations in CI", "test the audio pipeline", "check turn-taking and interruptions", or wants to check whole-conversation behavior before shipping. Covers text and audio mode and what each catches, running against a local or deployed agent, degraded-audio flags…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA. The repository describes itself as: A complete voice AI starter for LiveKit Agents with Python. The licence is MIT.

When your agent uses it

  • The user says run my simulations
  • Regression test my agent before deploying
  • Run the scenarios
  • Use lk agent simulate

Example prompts

  • “run my simulations”
  • “regression test my agent before deploying”
  • “run the scenarios”
  • “/running-livekit-simulations”

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. A bug in the agent. Fix the agent and re-run. Check instructions and tool descriptions first.
  2. A bad scenario. The expectation requires something the agent shouldn't do, or it's too vague
  3. A gap in the simulated world. The scenario passes but production failed, or the other way

What it can do on your machine

Read from SKILL.md and the folder at commit 76ddabb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Running Livekit Simulations loads about 1.7k tokens when it runs. Until then it costs about 211 tokens; SKILL.md has 880 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~211
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from livekit-examples/agent-starter-python at commit 76ddabb, republished under its MIT licence (© livekit-examples). 880 words, ~1,654 tokens.

Download SKILL.mdSave it as .claude/skills/running-livekit-simulations/SKILL.md (or your agent's skills folder).
name
running-livekit-simulations
description
Runs LiveKit agent simulations and acts on the results. Use when the user says "run my simulations", "regression test my agent before deploying", "run the scenarios", "use lk agent simulate", "did my agent pass", "why did this scenario fail", "run simulations in CI", "test the audio pipeline", "check turn-taking and interruptions", or wants to check whole-conversation behavior before shipping. Covers text and audio mode and what each catches, running against a local or deployed agent, degraded-audio flags, automating a pre-release run, and triaging failures with list, view and export. For writing the scenarios use writing-livekit-scenarios. Not the default for a bare "test my agent", which goes to debugging-livekit-agents. Use this skill when the user names simulations, scenarios, a run, CI, or shipping.
license
MIT
metadata.author
livekit

Running LiveKit simulations

A simulation plays a scenario against the real agent using an LLM-driven simulated user, then a judge grades the transcript. A unit test asserts on one turn. A simulation tells you whether a whole conversation reached the right outcome.

Read lk agent simulate --help before running. Subcommands and flags change, a wrong flag wastes a paid run, and this skill doesn't restate them. reading-livekit-docs has the rest.

When to reach for a simulation

Use simulations to regression-test long-horizon behavior before deploying to production: whether a multi-turn conversation reaches the right outcome when the caller backtracks, whether details gathered early survive to the end, whether the agent holds to its instructions under pressure, and whether it ended in the right state. For a single turn, use testing-livekit-agents; to poke at behavior while editing, use debugging-livekit-agents.

Running

Run from the agent's project directory. The mode is a subcommand:

bash
lk agent simulate text --scenarios scenarios.yaml    # see --help for the current flags

With no scenario file, the CLI generates scenarios from the agent's source. That uploads the code, and the CLI asks for confirmation first. Generation belongs to writing-livekit-scenarios.

By default the CLI starts the agent as a local worker, dispatches the scenarios to it, and stops it when the run ends. An option lets you grade an already-running agent by name instead. That needs a scenario file, since there's no local source to generate from.

Concurrency is limited per run and per project. The docs have the current limits.

Text or audio

Text is the default, and it's the right one. The simulated user exchanges text with the agent, so the run exercises the LLM, the tools and the conversation logic while the framework turns off STT, TTS and VAD. It's faster, cheaper and more deterministic. Use it for iteration and for anything automated.

Audio runs the same scenarios through the full speech pipeline. The simulated user speaks, listens and interrupts like a caller would, and the run scores what only speech exposes:

  • Turn-taking: starting to speak before the caller has finished, or leaving a caller who has finished waiting.
  • Interruption handling: yielding to a barge-in, and telling a brief acknowledgment apart from a new turn.
  • Transcription accuracy in both directions, scored separately for the things that matter: names, numbers, addresses, confirmation codes.
  • Perceived latency: what the caller heard, which differs from what the agent reports about itself. The gap between the two is what the user experiences.

Audio runs execute in real time, call the STT and TTS providers every turn, and are metered at a higher rate. Save them for a release candidate or a change that touches speech, turn-taking or interruption. Don't put them in a recurring job.

The audio subcommand has options to degrade the simulated caller's audio (noise, a poor microphone, packet loss). Use them to test what the agent does with speech it can't hear clearly. It should ask for a repeat instead of guessing. Combine them for a worst-case caller.

Show full SKILL.md (395 more words)Show less

Automating a pre-release run

All you need is a committed scenario file and a scheduled or release-branch job. The CLI prints plain output when it isn't attached to a terminal and exits non-zero when any scenario fails, so the job fails without extra wiring; the docs have a worked CI example to start from. Keep automated runs in text mode. Every scenario in the committed file has to pass or the job fails, so keep aspirational scenarios the agent doesn't pass yet in a separate file you run on demand.

Reading the results

A run prints a verdict per scenario and a dashboard link. The verdict tells you what happened; the transcript tells you why, so work from the transcript. The dashboard link is for the human. Your path is export: it prints a finished run, with each scenario's full chat context, as JSON — read a failing transcript from there, diff two runs, or archive a run as a build artifact. list finds the run id and also has machine-readable output; --help names the flags.

To triage a failure, decide which of these it is:

  1. A bug in the agent. Fix the agent and re-run. Check instructions and tool descriptions first. A failure that looks like bad reasoning is often a tool whose description never says when to use it.
  2. A bad scenario. The expectation requires something the agent shouldn't do, or it's too vague for a judge to decide consistently. Fix the scenario. A vague agent_expectations is the most common reason a verdict flips between runs.
  3. A gap in the simulated world. The scenario passes but production failed, or the other way round. Usually the agent reached different data, the derived instructions leave out the turn that caused the problem, or the failure only happens in audio and a text run can't see it.

After a fix, run the whole file, not only the scenario you were working on. A fix for one conversation often changes a neighbouring one.

Move repeat failures down the stack. A scenario that fails the same way every time is describing a turn-level bug. A unit test pins it more cheaply and catches it earlier. See testing-livekit-agents.

  • Authoring and organizing scenarios: writing-livekit-scenarios
  • Interactive debugging of a failure: debugging-livekit-agents
  • Cheaper per-commit coverage: testing-livekit-agents
  • Deploying the agent the run is grading: operating-livekit-agents
  • Flags, versions, changelogs: reading-livekit-docs

© livekit-examples, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/running-livekit-simulations of livekit-examples/agent-starter-python.

Open the folder on GitHubat commit 76ddabb

Used in 2 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in livekit-examples/agent-starter-python, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Running Livekit Simulations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Running Livekit Simulations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Running Livekit Simulations this skilllivekit-examples/agent-starter-python2641 repos~1.7kAutomated safety check: PassMIT
Codex Plugin QAcode-yeongyu/oh-my-openagent70k1 repos~1.9kAutomated safety check: PassCustom licence
Pester Failure AnalysisPowerShell/PowerShell56k—~5.1kAutomated safety check: PassMIT
Rust TDD Workflowrtk-ai/rtk83k—~753Automated safety check: NotesApache-2.0
Clawteam DevHKUDS/ClawTeam5.5k1 repos~1.1kAutomated safety check: PassMIT
Browser Testing with Chrome DevToolsaddyosmani/agent-skills103k4 repos~3.5kAutomated safety check: WarnMIT

Similar skills

  • Codex Plugin QA

    code-yeongyu/oh-my-openagent

    Tests the omo Codex plugin in an isolated CODEX_HOME with a local mock model, proving hooks fired through app-server notifications without touching ~/.codex.

    70k GitHub starsUsed in 1 repo~1.9k tokens
    Testing & QAAuto-check passed
  • Pester Failure Analysis

    PowerShell/PowerShell

    Investigates failing Pester tests in PowerShell CI jobs by following a six-step workflow from pull request status to documented fix recommendations.

    56k GitHub stars~5.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Enforces red-green-refactor for Rust work, with idiomatic test patterns, a naming convention and a pre-commit gate of cargo fmt, clippy and test.

    83k GitHub stars~753 tokensUpdated today
    Testing & QAAuto-check: notes
  • Clawteam Dev

    HKUDS/ClawTeam

    A skill your agent uses when working inside the ClawTeam repository itself: local development, debugging, reviewing, testing, validating multi-agent flows, or checking whether a code change actually…

    5.5k GitHub starsUsed in 1 repo~1.1k tokens
    Testing & QAAuto-check passed
  • Connects an agent to a real Chrome instance through the Chrome DevTools MCP server, so it can inspect the DOM, read console errors and profile performance directly.

    103k GitHub starsUsed in 4 repos~3.5k tokens
    Testing & QAAuto-check: warnings
  • Apple Container Test Runner

    RustPython/RustPython

    Runs RustPython tests inside a Linux container built with Apple's container CLI, so macOS users can compare Linux results with their local ones.

    22k GitHub stars~467 tokensUpdated today
    Testing & QAAuto-check passed

More from livekit-examples/agent-starter-python

  • Writing Livekit Scenarios

    livekit-examples/agent-starter-python

    Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them.

    264 GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Debugging Livekit Agents

    livekit-examples/agent-starter-python

    Drives a multi-turn conversation with a LiveKit agent running locally to see what it does.

    264 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Reading Livekit Docs

    livekit-examples/agent-starter-python

    Looks up current LiveKit facts (API signatures, CLI flags, config options, model and provider support, SDK changelogs, pricing) from the docs instead of answering from memory.

    264 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Building Livekit Agents

    livekit-examples/agent-starter-python

    Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud.

    264 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check: notes
  • Operating Livekit Agents

    livekit-examples/agent-starter-python

    Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside…

    264 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Testing Livekit Agents

    livekit-examples/agent-starter-python

    Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

    264 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed

Questions about Running Livekit Simulations

What does Running Livekit Simulations do?

Runs LiveKit agent simulations and acts on the results. An agent skill from livekit-examples/agent-starter-python. Running Livekit Simulations is an agent skill from livekit-examples/agent-starter-python. Runs LiveKit agent simulations and acts on the results.

When should I use Running Livekit Simulations?

Running Livekit Simulations fits situations like: the user says run my simulations; regression test my agent before deploying; run the scenarios; use lk agent simulate.

How do I install Running Livekit Simulations in Claude Code?

Run `npx skills add livekit-examples/agent-starter-python --skill running-livekit-simulations -a claude-code`. Or copy the skill folder (.agents/skills/running-livekit-simulations in livekit-examples/agent-starter-python) into .claude/skills/running-livekit-simulations in your project. Claude Code loads it when a task matches its description.

How do I install Running Livekit Simulations in Codex?

Run `npx skills add livekit-examples/agent-starter-python --skill running-livekit-simulations -a codex`. Or copy the skill folder (.agents/skills/running-livekit-simulations in livekit-examples/agent-starter-python) into .agents/skills/running-livekit-simulations in your project. Codex loads it when a task matches its description.

Can I use Running Livekit Simulations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add livekit-examples/agent-starter-python --skill running-livekit-simulations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/running-livekit-simulations, .gemini/skills/running-livekit-simulations, .github/skills/running-livekit-simulations and .opencode/skills/running-livekit-simulations in your project.

What does Running Livekit Simulations need to run?

SKILL.md names no scripts, command-line tools or credentials: Running Livekit Simulations is instructions for the agent only.

Does Running Livekit Simulations access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Running Livekit Simulations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Running Livekit Simulations use?

Running Livekit Simulations is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Running Livekit Simulations use?

About 1.7k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Running Livekit Simulations?

Skills that share tags, products or a category with Running Livekit Simulations: Codex Plugin QA (code-yeongyu/oh-my-openagent, 70k stars), Pester Failure Analysis (PowerShell/PowerShell, 56k stars), Rust TDD Workflow (rtk-ai/rtk, 83k stars) and Clawteam Dev (HKUDS/ClawTeam, 5.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Running Livekit Simulations?

livekit-examples (a GitHub organization) maintains it in livekit-examples/agent-starter-python, which has 264 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 6, 2026.

Source: livekit-examples/agent-starter-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.