Agent skill

Writing Livekit Scenarios

by livekit-examples in livekit-examples/agent-starter-python

Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them.

MITAuto-check passedTesting & QA

Install Writing Livekit Scenarios

skills CLI
$ npx skills add livekit-examples/agent-starter-python --skill writing-livekit-scenarios -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install livekit-examples/agent-starter-python writing-livekit-scenarios --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/livekit-examples/agent-starter-python.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/writing-livekit-scenarios .claude/skills/writing-livekit-scenarios && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
writing-livekit-scenarios
GitHub stars
264
Used in
1 other repo
Token cost
~2.5k tokens
SKILL.md length
1,411 words
Files
4 (incl. references)
Skills in repo
7
Repo updated
First seen
Licence
MIT

At a glance

Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them.

  • Works in 4 steps: Read every scenario and delete the ones… → Sharpen vague expectations. If the judge… → Add what generation misses. It drifts… → …
  • The user asks what should I test
  • SKILL.md covers Start from a generated baseline, Ask what to probe, Ground scenarios in what the… and Write instructions the…, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Writing Livekit Scenarios is an agent skill from livekit-examples/agent-starter-python. Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my simulations are flaky", "the scenario hits my real database", "seed state per scenario", "it passed but booked the wrong thing", "grade the final state", or wants to stress-test a flow before shipping. Covers generating a baseline with the…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/connecting-the-agent.md`, `references/risk-coverage.md` and `references/scenario-craft.md`).

It sits in Testing & QA, covering Load testing. The repository describes itself as: A complete voice AI starter for LiveKit Agents with Python. The licence is MIT.

When your agent uses it

  • The user asks what should I test
  • Generate simulation scenarios
  • Write scenarios for my agent
  • Add a scenario for X

Example prompts

  • “what should I test”
  • “generate simulation scenarios”
  • “write scenarios for my agent”
  • “/writing-livekit-scenarios”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Read every scenario and delete the ones that don't matter. A file you haven't read isn't a
  2. Sharpen vague expectations. If the judge can't decide an agent_expectations, its verdict
  3. Add what generation misses. It drifts toward happy paths. references/risk-coverage.md
  4. Add what the user is worried about. If they haven't said, ask what they want stress-tested.

What it can do on your machine

Read from SKILL.md and the folder at commit 76ddabb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Writing Livekit Scenarios loads about 2.5k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 219 tokens; SKILL.md has 1,411 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~219
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from livekit-examples/agent-starter-python at commit 76ddabb, republished under its MIT licence (© livekit-examples). 1,411 words, ~2,510 tokens.

Download SKILL.mdSave it as .claude/skills/writing-livekit-scenarios/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
writing-livekit-scenarios
description
Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my simulations are flaky", "the scenario hits my real database", "seed state per scenario", "it passed but booked the wrong thing", "grade the final state", or wants to stress-test a flow before shipping. Covers generating a baseline with the LiveKit scenario generator and refining it, writing the cases generation misses, phrasing simulated-user instructions so they steer reliably, splitting scenarios into sets across files, and the agent-side code that seeds deterministic state from a scenario and fails a run on its final state. To run them use running-livekit-simulations.
license
MIT
metadata.author
livekit

Writing simulation scenarios

A scenario is a simulated user's instructions (who they are, what they want) plus agent_expectations (what counts as success). A simulation plays the scenario against the real agent, and an LLM judge grades the transcript.

Runs are disposable. The scenario file is what lasts: it gets reviewed in diffs, re-run for years, and it's what catches a prompt edit that breaks something without anyone noticing.

Before writing a file, look up the exact schema and CLI flags with reading-livekit-docs. Field names and commands change, so this skill doesn't restate them. Conceptually, a file is a named group of scenarios. Each scenario has the simulated user's instructions, the pass criteria, and optionally tags for grouping and per-scenario data the agent can read at runtime. If the CLI offers to add stable per-scenario ids to a file, accept and commit them.

Start from a generated baseline

Don't start from an empty file. LiveKit's generator reads the agent's source and produces a decent first pass much faster than you'd write ten scenarios by hand:

bash
lk agent simulate text -n 10    # confirm the exact flags with --help

Generating from source uploads the code to LiveKit Cloud so the generator can read it, which is why the CLI asks for confirmation first. Tell the user that before running it and let them decline. Some code can't leave the machine, and in that case you write the scenarios by hand. A flag skips the prompt for non-interactive runs.

When the run finishes, the CLI either tells you where it saved the generated scenarios or offers to save them into the project. Either way, get that file into the repo as the starting point.

You can't steer the generator. It infers intent from the code, so it writes plausible conversations instead of the ones this agent's users have, and it doesn't know what the user is worried about. Treat its output as a draft:

  1. Read every scenario and delete the ones that don't matter. A file you haven't read isn't a test suite.
  2. Sharpen vague expectations. If the judge can't decide an agent_expectations, its verdict flips between runs. This is the biggest cause of flaky scenarios.
  3. Add what generation misses. It drifts toward happy paths. references/risk-coverage.md explains how to turn the agent's constraints into a checklist and give each item a scenario.
  4. Add what the user is worried about. If they haven't said, ask what they want stress-tested. That input is why a suite a person steers is better than a generated one.

Ask what to probe

The user knows which flow keeps breaking, which customer complained, and which change they're nervous about. Ask, and weight the suite toward that area with more and deeper scenarios.

Focus adds coverage. It never removes coverage of the agent's hard limits. If the user has no preference, generate broadly and tell them that's what you did.

Ground scenarios in what the agent can do

Read the agent's code locally with your normal tools before writing anything. Look for what a user can ask for (capabilities), where requests get blocked (constraints: required steps, unavailable items by name, caps, eligibility), and what the agent must refuse.

Constraints matter most. A scenario that asks for something the agent can't do is only valid if the expectation is that the agent says so. Written the other way, the test fails when the agent behaves correctly.

Scope is the agent reachable from the session entrypoint, plus agents it hands off to and tasks it awaits. Ignore other classes in the directory, unused imports, and example files.

Write instructions the simulated user can follow

There are two shapes, and which one you use depends on what the scenario is for. Details and examples are in references/scenario-craft.md.

  • Point-form direction for deterministic flows you'll grade on end state. A structured template (persona, opening line, facts revealed only when asked, steps in order, conditional reactions) keeps the persona model on track and makes the run repeatable.
  • A persona paragraph plus goals for open-ended and adversarial scenarios graded only on the conversation, where you want the simulated user to improvise.

In both shapes, goals are requests to the agent, never the agent's own actions. Use real values from the agent's domain and assume no prior state. Vary persona, mood and difficulty across the suite so it isn't ten copies of the same cooperative caller.

Split scenarios into sets, one file per set

The run command takes one scenario file, so a file is a run. Split along the lines you want to run separately, which usually isn't topic.

The most important axis is how the set is graded:

  • Scenarios that drive concrete tool flows to a deterministic end state, graded on final state and the conversation (see "Make the agent consume the scenario" below).
  • Open-ended and adversarial scenarios graded on agent_expectations alone.

These need different instructions shapes, different agent wiring, and often a different run cadence, so they go in different files. The second axis is cadence: a small set to run before a merge, and the full set before a release.

Give each file a name that describes the set, since it labels the run. Inside a file, use tags to slice. Tag feature so a failure points at the part of the agent that owns it, and add whatever else you filter by (channel, difficulty). Put a header comment on the file recording the set's assumptions: pinned dates, required environment variables, and what its expectations depend on.

Show full SKILL.md (507 more words)Show less

Keep scenarios reproducible

A scenario should give the same result months from now as it does today.

  • Write absolute dates and pin the agent's clock (usually with an environment variable) so availability and expectations always line up. Relative dates rot.
  • Don't let scenarios hit real backends. Seed deterministic state from the scenario instead (see "Make the agent consume the scenario" below).
  • Keep ids stable where the CLI supports them. Labels and instructions change, and the id is what ties a scenario's runs together over time. A scenario rewritten from scratch is a new scenario and gets a new id.

Make the agent consume the scenario

A scenario file can't fix two problems by itself. A scenario that reaches a real backend grades differently every run, and a judge that only reads the transcript will pass a run that booked the wrong room. Both are fixed in the agent's code, in one branch at the top of the entrypoint:

  1. Detect a simulation from the job context. In production the check comes back empty and nothing else changes.
  2. Seed state from the scenario's per-scenario data (a fake calendar, an in-memory database) that the agent's real tools read through unchanged.
  3. Mock tools for the session's lifetime, so tool selection is still under test but execution is deterministic.
  4. Grade the final state in the end-of-simulation callback. Compare what the agent ended with against what the scenario expected, and fail the run if they differ. Your check can fail a run the judge passed. It can't rescue one the judge failed.

Only tool-flow scenarios need this. Open-ended and adversarial ones are graded on the conversation alone. references/connecting-the-agent.md covers it in full, including how to confirm the wiring took effect. Look up the current API names with reading-livekit-docs.

Grow the suite from real failures

The most valuable scenarios describe something that went wrong. Once the agent is live, derive them from recorded sessions instead of inventing more. Where the CLI supports it, a subcommand derives a scenario from a recorded session (check lk agent simulate --help). A derived scenario describes one call, so widen it to the class of calls it represents and sharpen its expectation into the rule you want enforced.

Every production bug you fix should get a permanent scenario in the file.

Don't write bad tests

The judge grades the agent against agent_expectations, so a careless expectation punishes correct behavior:

  • For guardrails and negative cases, the pass is that the agent refuses, escalates, or declines to invent data. Write the expectation that way.
  • Never write an expectation the agent can only meet by misbehaving, such as stating data it can't know or giving specific medical, legal or financial directives. If passing requires misbehavior, the scenario is wrong.
  • Judge by outcome, from the user's point of view. Don't prescribe wording.

References

  • references/risk-coverage.md — turning constraints into a checklist and guaranteeing coverage
  • references/scenario-craft.md — instructions shapes, persona variety, audio-specific scenarios
  • references/connecting-the-agent.md — the agent-side code: detect, seed, mock, grade final state
  • Running them: running-livekit-simulations
  • Schema and CLI facts: reading-livekit-docs

© livekit-examples, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .agents/skills/writing-livekit-scenarios of livekit-examples/agent-starter-python.

  • SKILL.md
  • references/connecting-the-agent.md
  • references/risk-coverage.md
  • references/scenario-craft.md

Open the folder on GitHubat commit 76ddabb

Used in 2 other repositories

We found 4 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in livekit-examples/agent-starter-python, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Writing Livekit Scenarios next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Writing Livekit Scenarios compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Writing Livekit Scenarios this skilllivekit-examples/agent-starter-python2641 repos~2.5kAutomated safety check: PassMIT
Go Testingcxuu/golang-skills1731 repos~1.3kAutomated safety check: PassApache-2.0
Goalcraftgrp06/goalcraft102—~3.8kAutomated safety check: PassMIT
Thinking Partnermattnowdev/thinking-partner206—~4.4kAutomated safety check: PassMIT
Visionkunchenguid/vision331—~2.9kAutomated safety check: PassMIT
Volt Load TestingowenHochwald/volt141—~1.2kAutomated safety check: PassMPL-2.0

Similar skills

  • Go Testing

    cxuu/golang-skills

    A skill your agent uses when writing, reviewing, or improving Go test code — including table-driven tests, subtests, parallel tests, test helpers, test doubles, and assertions with cmp.Diff.

    173 GitHub starsUsed in 1 repo~1.3k tokens
    Testing & QAAuto-check passed
  • Goalcraft

    grp06/goalcraft

    Turn a rough draft, vague ambition, or messy task brief into a powerful Codex /goal objective for persistent, evidence-checked work.

    102 GitHub stars~3.8k tokensUpdated 4 mo ago
    Testing & QAAuto-check passed
  • Thinking Partner

    mattnowdev/thinking-partner

    A deterministic thinking partner that challenges assumptions and applies mental models to sharpen decisions, solve problems, and think more clearly.

    206 GitHub stars~4.4k tokensUpdated 6 mo ago
    Testing & QAAuto-check passed
  • Vision

    kunchenguid/vision

    Draft and stress-test a VISION.md for a repository, then iterate with the author on an interactive review board until approved.

    331 GitHub stars~2.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Volt Load Testing

    owenHochwald/volt

    Safely exercise and evaluate HTTP APIs with the Volt CLI, including authenticated requests, JSON bodies, staged load, machine-readable results, performance baselines, and before/after comparisons.

    141 GitHub stars~1.2k tokensUpdated 2 mo ago
    Testing & QAAuto-check passed
  • 5 Persona Advisory Board

    harryvondiesel-web/5-persona-advisory-board

    Run a 5 Persona Advisory Board review, board review, strategic decision stress-test, offer critique, risk check, or pricing/timing/positioning decision review.

    133 GitHub stars~2.8k tokensUpdated 3 days ago
    Testing & QAAuto-check passed

More from livekit-examples/agent-starter-python

  • Debugging Livekit Agents

    livekit-examples/agent-starter-python

    Drives a multi-turn conversation with a LiveKit agent running locally to see what it does.

    264 GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Reading Livekit Docs

    livekit-examples/agent-starter-python

    Looks up current LiveKit facts (API signatures, CLI flags, config options, model and provider support, SDK changelogs, pricing) from the docs instead of answering from memory.

    264 GitHub starsUsed in 1 repo~1.1k tokens
    Auto-check passed
  • Running Livekit Simulations

    livekit-examples/agent-starter-python

    Runs LiveKit agent simulations and acts on the results. An agent skill from livekit-examples/agent-starter-python.

    264 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Building Livekit Agents

    livekit-examples/agent-starter-python

    Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud.

    264 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check: notes
  • Operating Livekit Agents

    livekit-examples/agent-starter-python

    Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside…

    264 GitHub starsUsed in 1 repo~2.4k tokens
    Auto-check passed
  • Testing Livekit Agents

    livekit-examples/agent-starter-python

    Writes turn-level tests for a LiveKit agent in the user's normal test suite: pytest (Python) or Vitest (Node.js).

    264 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed

Categories

Questions about Writing Livekit Scenarios

What does Writing Livekit Scenarios do?

Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Writing Livekit Scenarios is an agent skill from livekit-examples/agent-starter-python. Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them.

When should I use Writing Livekit Scenarios?

Writing Livekit Scenarios fits situations like: the user asks what should I test; generate simulation scenarios; write scenarios for my agent; add a scenario for X.

How do I install Writing Livekit Scenarios in Claude Code?

Run `npx skills add livekit-examples/agent-starter-python --skill writing-livekit-scenarios -a claude-code`. Or copy the skill folder (.agents/skills/writing-livekit-scenarios in livekit-examples/agent-starter-python) into .claude/skills/writing-livekit-scenarios in your project. Claude Code loads it when a task matches its description.

How do I install Writing Livekit Scenarios in Codex?

Run `npx skills add livekit-examples/agent-starter-python --skill writing-livekit-scenarios -a codex`. Or copy the skill folder (.agents/skills/writing-livekit-scenarios in livekit-examples/agent-starter-python) into .agents/skills/writing-livekit-scenarios in your project. Codex loads it when a task matches its description.

Can I use Writing Livekit Scenarios in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add livekit-examples/agent-starter-python --skill writing-livekit-scenarios -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/writing-livekit-scenarios, .gemini/skills/writing-livekit-scenarios, .github/skills/writing-livekit-scenarios and .opencode/skills/writing-livekit-scenarios in your project.

What does Writing Livekit Scenarios need to run?

SKILL.md names no scripts, command-line tools or credentials: Writing Livekit Scenarios is instructions for the agent only.

Does Writing Livekit Scenarios access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Writing Livekit Scenarios safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Writing Livekit Scenarios use?

Writing Livekit Scenarios is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Writing Livekit Scenarios use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.6k tokens, read only when the agent opens those files.

What are the alternatives to Writing Livekit Scenarios?

Skills that share tags, products or a category with Writing Livekit Scenarios: Go Testing (cxuu/golang-skills, 173 stars), Goalcraft (grp06/goalcraft, 102 stars), Thinking Partner (mattnowdev/thinking-partner, 206 stars) and Vision (kunchenguid/vision, 331 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Writing Livekit Scenarios?

livekit-examples (a GitHub organization) maintains it in livekit-examples/agent-starter-python, which has 264 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: livekit-examples/agent-starter-python on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.