Agent skill

Talking Avatar Voice Chat App

by buildfastwithai in buildfastwithai/gen-ai-experiments

Builds a realtime voice-chat app around a talking character portrait made from your photo or a text description, with mouth sprites driven by the audio.

MITAuto-check passedAI & LLM Engineering

Install Talking Avatar Voice Chat App

skills CLI
$ npx skills add buildfastwithai/gen-ai-experiments --skill talking-avatar -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install buildfastwithai/gen-ai-experiments talking-avatar --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/buildfastwithai/gen-ai-experiments.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/talking-avatar .claude/skills/talking-avatar && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
talking-avatar
GitHub stars
785
Token cost
~1.7k tokens
SKILL.md length
832 words
Files
14 (incl. scripts, references, assets)
Skills in repo
13
Repo updated
First seen
Licence
MIT

At a glance

Builds a realtime voice-chat app around a talking character portrait made from your photo or a text description, with mouth sprites driven by the audio.

  • Works in 3 steps: Photograph supplied: inspect it with… → Character described: normalize the… → Establish the character name, concise…
  • Creating a talking photo avatar you can speak with in the browser
  • SKILL.md covers Coordinate the required skills, Choose the input path, Build the asset set and Build or integrate the app, plus 4 more sections
  • Calls python3

What it does

The goal is one focused voice conversation with a fixed character portrait. From a supplied photo, treated as an identity reference, or a described character, the agent generates a canonical front-facing portrait, then three identity-preserving mouth edits named soft, round and open, rejecting any frame where more than the lips changed. The mouth region is cropped identically across the three frames so the app can swap sprites as the remote audio plays.

The app is built on OpenAI Realtime in Vite or Next.js, starting from bundled templates that include a session route and a test file. Image generation is handed to a separate image skill, and the agent is told to check OpenAI's current documentation for the Realtime endpoint, session schema, model and voices instead of trusting the skill's own examples. It covers bring-your-own API key intake, testing and deployment, and adds no backgrounds, wardrobes or extra characters by default.

When your agent uses it

  • Creating a talking photo avatar you can speak with in the browser
  • Building a character chatbot or roleplay voice companion
  • Adding audio-driven mouth animation to a generated character portrait

Example prompts

  • “Turn the photo at ./me.jpg into a talking avatar I can chat with by voice.”
  • “Build a pirate captain voice companion with an illustrated look and a deep voice.”
  • “Add lip sync to this Next.js voice app using three mouth sprites.”

Requirements

  • An OpenAI API key supplied by you
  • An image generation skill or tool for the portrait and mouth frames

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Photograph supplied: inspect it with view_image, treat it as an identity reference, and preserve face shape, skin tone, hair, glasses…
  2. Character described: normalize the description into a concrete visual spec. Ask only for a missing detail that materially changes…
  3. Establish the character name, concise persona, app name, language, and preferred voice. Infer reasonable defaults when safe.

What it can do on your machine

Read from SKILL.md and the folder at commit 7b62043. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Talking Avatar Voice Chat App loads about 1.7k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 119 tokens; SKILL.md has 832 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~119
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from buildfastwithai/gen-ai-experiments at commit 7b62043, republished under its MIT licence (© buildfastwithai). 832 words, ~1,747 tokens.

Download SKILL.mdSave it as .claude/skills/talking-avatar/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.
name
talking-avatar
description
Build or transform lightweight realtime voice-chat apps with a talking character avatar from either a user-supplied photograph or a text description. Use when Codex is asked to create a talking avatar, photo avatar, character chatbot, roleplaying voice companion, mouth-sprite animation, or reusable OpenAI Realtime avatar app in Vite/Next.js, including BYOK API-key intake, identity-consistent image generation, audio-driven lip sync, testing, and deployment.

Talking Avatar

Create one focused voice conversation with a fixed character portrait whose mouth poses follow the actual remote audio. Accept either a photograph or a character description, generate a canonical portrait plus small mouth sprites, and build the app around OpenAI Realtime.

Coordinate the required skills

  • Use imagegen for every canonical portrait or mouth-frame generation/edit. Read its full instructions before generating.
  • Use openai-docs and, when applicable, openai-platform-api-key to verify the current Realtime WebRTC endpoint, session schema, model, voices, and credential flow. Never rely on this skill's model examples as current truth.
  • Use the available frontend design skill for interface work.
  • If .openai/hosting.json exists, follow sites-building and sites-hosting through deployment.
  • Do not add backgrounds, wardrobes, role modes, or multiple characters unless the user explicitly asks. The default product is one avatar and one conversation.

Choose the input path

  1. Photograph supplied: inspect it with view_image, treat it as an identity reference, and preserve face shape, skin tone, hair, glasses, facial hair, and distinguishing features.
  2. Character described: normalize the description into a concrete visual spec. Ask only for a missing detail that materially changes identity, such as photorealistic versus illustrated. Otherwise proceed with tasteful defaults.
  3. Establish the character name, concise persona, app name, language, and preferred voice. Infer reasonable defaults when safe.

Read references/image-pipeline.md before creating or editing the avatar assets.

Build the asset set

  1. Generate one canonical, front-facing, evenly lit, closed-mouth portrait. Lock its canvas, crop, head position, gaze, pose, clothing, background, and lighting.
  2. From that canonical image, generate three separate identity-preserving edits:
    • soft: lips slightly parted;
    • round: a natural rounded vowel;
    • open: a natural medium-open vowel.
  3. Repeat the invariants in every edit prompt. Change only the immediate lips and inner mouth. Reject frames with moved eyes, jaw, glasses, head, beard, clothing, or lighting.
  4. Crop the smallest practical identical rectangle from all three speech frames. Keep enough surrounding skin for feathering, but do not ship full-frame speech images.
  5. Save exactly these project assets by default:
    • public/avatar/avatar-base.jpg
    • public/avatar/mouth-soft.png
    • public/avatar/mouth-round.png
    • public/avatar/mouth-open.png
  6. Run scripts/validate_avatar_assets.py --dir <project>/public/avatar and fix all failures.

Never animate the complete portrait. Only replace the mouth patch.

Build or integrate the app

Prefer an existing working stack. For a new project, use Next.js or vinext because the user's API key must cross a server route; use plain Vite only when a safe backend already exists.

For a compatible new Next/vinext project, scaffold the proven core:

bash
python3 <skill-dir>/scripts/scaffold_app.py \
  --target <project-dir> \
  --character-name "Mira" \
  --app-name "Mira Live" \
  --persona "a concise, curious field researcher" \
  --model "<verified-realtime-model>" \
  --transcription-model "<verified-transcription-model>" \
  --voice "<verified-realtime-voice>"

The target must already contain the framework configuration and dependencies. Install framer-motion and @phosphor-icons/react when missing. Use --force only for a newly initialized project after inspecting the files it will replace. For an existing product, copy or adapt individual files from assets/starter/ instead of overwriting its architecture.

Read references/app-contract.md before changing the Realtime route, API-key flow, or event handling.

Show full SKILL.md (370 more words)Show less

Implement natural lip sync

Read references/realtime-lipsync.md and preserve its behavioral contract:

  • Analyze the remote output MediaStream, not microphone input and not Realtime event cadence.
  • Sample audio every animation frame without updating React state.
  • Update the visible mouth pose no faster than about every 96 ms.
  • Smooth attack and decay, normalize against a rolling peak, and close after short silence.
  • Move through adjacent poses (closed → soft → round → open) instead of jumping.
  • Set a DOM data-mouth attribute and let CSS reveal exactly one patch.
  • Keep the portrait pixels completely static.

Do not use a canned infinite mouth animation. response.output_audio.delta identifies speaking state but does not provide lip timing.

Preserve the BYOK security boundary

  • Ask the user to enter the OpenAI API key inside the app.
  • Keep the key in component memory only; never use local storage, session storage, IndexedDB, analytics, logs, or source files.
  • Send it once to the same-origin session route over HTTPS, clear client state after negotiation, and return Cache-Control: no-store.
  • Never expose a long-lived key directly to the Realtime WebRTC peer connection.
  • Map upstream authentication and rate-limit failures to short user-facing errors without echoing response bodies or secrets.

Validate before handoff

  1. Run the production build.
  2. Run lint and the starter regression test when available.
  3. Confirm the source contains an AnalyserNode, requestAnimationFrame, a 90–110 ms pose interval, adjacent-pose stepping, and no CSS @keyframes mouth-* loop.
  4. Confirm all mouth patches have identical dimensions and align over the canonical mouth at desktop and mobile sizes.
  5. Confirm only the mouth changes during speech; blinking, head bobbing, pose drift, and full-frame swaps are failures.
  6. Confirm mic mute, speaker mute, interruption, transcript deltas, typed messages, disconnect cleanup, audio-context cleanup, empty state, connection errors, and invalid-key errors.
  7. If the user requests visual browser QA, test a live voice turn and tune timing from observed speech. Otherwise do not claim that browser-level lip sync was visually verified.
  8. Deploy when requested or when the active hosting skill requires it, then return the live URL.

Deliverables

Return the runnable app, the final avatar asset paths, the generation prompt set, whether photo or description mode was used, the Realtime model/voice verified for that run, validation results, and the deployed URL when applicable.

© buildfastwithai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 13 other files (scripts, references, assets) in skills/talking-avatar of buildfastwithai/gen-ai-experiments.

  • SKILL.md
  • agents/openai.yaml
  • assets/starter/app/TalkingAvatarApp.tsx.tmpl
  • assets/starter/app/api/realtime/session/route.ts.tmpl
  • assets/starter/app/globals.css.tmpl
  • assets/starter/app/layout.tsx.tmpl
  • assets/starter/app/page.tsx.tmpl
  • assets/starter/tests/talking-avatar.test.mjs.tmpl
  • references/app-contract.md
  • references/image-pipeline.md
  • references/realtime-lipsync.md
  • scripts
  • … and 2 more

Open the folder on GitHubat commit 7b62043

Compare with similar skills

Talking Avatar Voice Chat App next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Talking Avatar Voice Chat App compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Talking Avatar Voice Chat App this skillbuildfastwithai/gen-ai-experiments785—~1.7kAutomated safety check: PassMIT
Azure AI Openai Dotnetmicrosoft/skills3.1k5 repos~3.4kAutomated safety check: PassMIT
Mem0 Provider for Vercel AI SDKmem0ai/mem067k—~2.3kAutomated safety check: PassApache-2.0
Local AI Useamd/skills406—~5kAutomated safety check: NotesMIT
Gpt Imagenikships/droidproxy122—~2.5kAutomated safety check: PassMIT
Fastllm Gatewayazrtydxb/Fastllm-proxy108—~926Automated safety check: PassApache-2.0

Similar skills

  • Azure AI Openai Dotnet

    microsoft/skills

    Official

    Azure OpenAI SDK for .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Adds persistent memory to Vercel AI SDK apps with the Mem0 provider, using a wrapped model or standalone retrieve and store utilities.

    67k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    406 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Gpt Image

    nikships/droidproxy

    Generate or edit images via GPT Image 2.5 Flare or Sunburst through DroidProxy Codex OAuth (no OPENAIAPIKEY).

    122 GitHub stars~2.5k tokensUpdated today
    Media & CreativeAuto-check passed
  • Fastllm Gateway

    azrtydxb/Fastllm-proxy

    Send inference requests through the FastLLM OpenAI-compatible gateway — chat completions, completions, embeddings, rerank, score, responses, moderations, audio speech and transcription, image…

    108 GitHub stars~926 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Integrates local AI capabilities into applications using Embeddable Lemonade.

    406 GitHub stars~6k tokensUpdated today
    Media & CreativeAuto-check passed

More from buildfastwithai/gen-ai-experiments

All 13 skills in this repo
  • Landing Page Generator

    buildfastwithai/gen-ai-experiments

    Builds a single self-contained HTML landing page from a template with copy frameworks, four themes and SEO meta, then runs bundled audit scripts for conversion and speed.

    785 GitHub stars~1.5k tokensUpdated 17 days ago
    Auto-check passed
  • LaunchAudit

    buildfastwithai/gen-ai-experiments

    Audits a startup, app or landing page from a URL, localhost, repository, screenshots or copy, then gives a launch-readiness verdict, prioritized fixes and an HTML report.

    785 GitHub stars~1.8k tokensUpdated 17 days ago
    Auto-check passed
  • Startupblueprint

    buildfastwithai/gen-ai-experiments

    Turn a startup, SaaS, app, developer tool, website, repository, or product idea into an evidence-backed business plan, monetization strategy, pricing architecture, editable 12-month financial model…

    785 GitHub stars~2.2k tokensUpdated 17 days ago
    Auto-check passed
  • Eazyr Repo Onboarding Kit

    buildfastwithai/gen-ai-experiments

    Creates a QUICKSTART, doctor scripts and a documented .env.example for an unfamiliar repo from facts scanned in the repo, and audits how easy it is to onboard.

    785 GitHub stars~1.4k tokensUpdated 17 days ago
    Auto-check: notes
  • Mutation Test Strength Audit

    buildfastwithai/gen-ai-experiments

    Measures how well a Python pytest suite catches behavior changes through diff-scoped mutation testing, then proposes and verifies tests for surviving mutants.

    785 GitHub stars~641 tokensUpdated 17 days ago
    Auto-check passed
  • Customer Finder

    buildfastwithai/gen-ai-experiments

    Find and qualify evidence-backed potential first customers, early adopters, design partners, or beta users for a startup using recent public pain and buying signals.

    785 GitHub stars~1.4k tokensUpdated 17 days ago
    Auto-check passed

Questions about Talking Avatar Voice Chat App

What does Talking Avatar Voice Chat App do?

Builds a realtime voice-chat app around a talking character portrait made from your photo or a text description, with mouth sprites driven by the audio. The goal is one focused voice conversation with a fixed character portrait. From a supplied photo, treated as an identity reference, or a described character, the agent generates a canonical front-facing portrait, then three identity-preserving mouth edits named soft, round and open, rejecting any frame where more than the lips changed.

When should I use Talking Avatar Voice Chat App?

Talking Avatar Voice Chat App fits situations like: creating a talking photo avatar you can speak with in the browser; building a character chatbot or roleplay voice companion; adding audio-driven mouth animation to a generated character portrait.

How do I install Talking Avatar Voice Chat App in Claude Code?

Run `npx skills add buildfastwithai/gen-ai-experiments --skill talking-avatar -a claude-code`. Or copy the skill folder (skills/talking-avatar in buildfastwithai/gen-ai-experiments) into .claude/skills/talking-avatar in your project. Claude Code loads it when a task matches its description.

How do I install Talking Avatar Voice Chat App in Codex?

Run `npx skills add buildfastwithai/gen-ai-experiments --skill talking-avatar -a codex`. Or copy the skill folder (skills/talking-avatar in buildfastwithai/gen-ai-experiments) into .agents/skills/talking-avatar in your project. Codex loads it when a task matches its description.

Can I use Talking Avatar Voice Chat App in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add buildfastwithai/gen-ai-experiments --skill talking-avatar -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/talking-avatar, .gemini/skills/talking-avatar, .github/skills/talking-avatar and .opencode/skills/talking-avatar in your project.

What does Talking Avatar Voice Chat App need to run?

Going by SKILL.md and its folder, Talking Avatar Voice Chat App needs the command-line tools its instructions call (python3). Our summary lists: An OpenAI API key supplied by you; An image generation skill or tool for the portrait and mouth frames.

Does Talking Avatar Voice Chat App access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Talking Avatar Voice Chat App safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Talking Avatar Voice Chat App use?

Talking Avatar Voice Chat App is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Talking Avatar Voice Chat App use?

About 1.7k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Talking Avatar Voice Chat App?

Skills that share tags, products or a category with Talking Avatar Voice Chat App: Azure AI Openai Dotnet (microsoft/skills, 3.1k stars), Mem0 Provider for Vercel AI SDK (mem0ai/mem0, 67k stars), Local AI Use (amd/skills, 406 stars) and Gpt Image (nikships/droidproxy, 122 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Talking Avatar Voice Chat App?

buildfastwithai (a GitHub organization) maintains it in buildfastwithai/gen-ai-experiments, which has 785 GitHub stars. The repository holds 13 skills in this directory. The repository was last updated on September 22, 2026.

Source: buildfastwithai/gen-ai-experiments on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.