Agent skill

Voice Interaction

by Owl-Listener in Owl-Listener/inclusive-design-skills

Design voice interactions and speech interfaces that work for people with diverse speech patterns, accents, and communication styles.

MITAuto-check passedAI & LLM Engineering

Install Voice Interaction

skills CLI
$ npx skills add Owl-Listener/inclusive-design-skills --skill voice-interaction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Owl-Listener/inclusive-design-skills voice-interaction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Owl-Listener/inclusive-design-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/inclusive-interaction/skills/voice-interaction .claude/skills/voice-interaction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice-interaction
GitHub stars
104
Token cost
~730 tokens
SKILL.md length
341 words
Files
1
Skills in repo
55
Repo updated
First seen
Licence
MIT

At a glance

Design voice interactions and speech interfaces that work for people with diverse speech patterns, accents, and communication styles.

  • Works in 5 steps: Can every voice-activated feature also… → Does the system handle varied speech… → Is there clear feedback showing what the… → …
  • Designing voice commands
  • SKILL.md covers Who This Is For, Core Principles, Design Patterns and Assessment Questions
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Voice Interaction is an agent skill from Owl-Listener/inclusive-design-skills. Design voice interactions and speech interfaces that work for people with diverse speech patterns, accents, and communication styles. Use when designing voice commands, voice search, dictation, voice assistants, or any interface that accepts speech input. Triggers on: voice, speech, dictation, voice command, voice search, speech recognition, accent, stutter, speech disability, non-verbal, AAC, voice assistant, talk to type.

Its SKILL.md is about 730 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Speech recognition and synthesis and Transcription. The repository describes itself as: Inclusive design skills for AI coding agents — from cognitive accessibility to adaptive interfaces, inclusive research, and accessibility decision-making. The licence is MIT.

When your agent uses it

  • Designing voice commands
  • Voice assistants
  • Any interface that accepts speech input
  • Speech recognition

Example prompts

  • “/voice-interaction”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Can every voice-activated feature also be completed without voice?
  2. Does the system handle varied speech patterns without frustration?
  3. Is there clear feedback showing what the system heard?
  4. Can users easily correct misrecognition?
  5. Is it obvious when the system is listening?

What it can do on your machine

Read from SKILL.md and the folder at commit 6e0740f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice Interaction loads about 730 tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 341 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~730

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Owl-Listener/inclusive-design-skills at commit 6e0740f, republished under its MIT licence (© Owl-Listener). 341 words, ~730 tokens.

Download SKILL.mdSave it as .claude/skills/voice-interaction/SKILL.md (or your agent's skills folder).
name
voice-interaction
description
Design voice interactions and speech interfaces that work for people with diverse speech patterns, accents, and communication styles. Use when designing voice commands, voice search, dictation, voice assistants, or any interface that accepts speech input. Triggers on: voice, speech, dictation, voice command, voice search, speech recognition, accent, stutter, speech disability, non-verbal, AAC, voice assistant, talk to type.

Voice Interaction Design

Design voice interfaces that work for the full range of human speech — including accents, speech disabilities, non-native speakers, and people in noisy or quiet environments.

Who This Is For

  • People with motor disabilities who use voice as primary input
  • People who stutter, have dysarthria, or other speech differences
  • Non-native speakers with varied accents
  • People in environments where typing is impractical
  • Anyone who prefers voice to typing for certain tasks

Core Principles

Voice Should Never Be the Only Option
  • Every voice interaction must have a text/touch/keyboard alternative
  • Voice is an accelerator, not a gatekeeper
  • Don't require voice for identity verification or critical actions unless an alternative exists
Design for Speech Variation
  • Support varied pacing — don't cut off slow speakers
  • Allow generous silence before timing out
  • Don't penalise repetition, filler words, or self-correction
  • Support multiple phrasings for the same intent ("go back", "previous page", "take me back", "undo")
Feedback Must Be Clear
  • Confirm what the system heard (visual transcript)
  • Make it easy to correct misrecognition
  • Show when the system is listening vs. processing vs. waiting
  • Never execute a destructive action on voice alone without confirmation

Design Patterns

Flexible Recognition
  • Accept multiple ways to say the same command
  • Don't require exact phrasing — intent matters more than syntax
  • Support "did you mean?" clarification for ambiguous input
  • Allow users to spell out words the system doesn't recognise
Graceful Failure
  • When recognition fails: show what was heard and offer correction
  • Never respond with just "I didn't understand" — offer alternatives
  • Provide a "type instead" option at every failure point
  • After repeated failures: proactively suggest switching to text input
Privacy and Control
  • Clear visual/audio indicator when microphone is active
  • Easy one-action mute/stop listening
  • Don't record or transmit audio without explicit consent
  • Allow users to review and delete voice data

Assessment Questions

  1. Can every voice-activated feature also be completed without voice?
  2. Does the system handle varied speech patterns without frustration?
  3. Is there clear feedback showing what the system heard?
  4. Can users easily correct misrecognition?
  5. Is it obvious when the system is listening?

© Owl-Listener, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in inclusive-interaction/skills/voice-interaction of Owl-Listener/inclusive-design-skills.

Open the folder on GitHubat commit 6e0740f

Compare with similar skills

Voice Interaction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice Interaction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice Interaction this skillOwl-Listener/inclusive-design-skills104—~730Automated safety check: PassMIT
Yichen Asrmcncarl/yichen-skills4.4k—~780Automated safety check: PassCustom licence
Youtube FetcherJimmySadek/youtube-fetcher-to-markdown485—~3.1kAutomated safety check: PassMIT
Volcengine Asrysyecust/lecture-to-notes273—~783Automated safety check: PassCustom licence
Openai Whisperhuangruiteng/CS-Notes4k18 repos~228Automated safety check: PassMIT
Ax Audiodosco/aithy107—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Yichen Asr

    mcncarl/yichen-skills

    逸尘自用的统一音视频转写入口,在 StepFun Step ASR 与火山引擎豆包 ASR 之间按输出需求、安全边界和可用状态路由。用于本地音频或视频的纯文本转写、时间戳、SRT 字幕、口播粗剪,以及转写前体检;用户明确指定服务商时不得静默切换。Use when a local audio or video file needs transcription and the correct…

    4.4k GitHub stars~780 tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Youtube Fetcher

    JimmySadek/youtube-fetcher-to-markdown

    Retrieve transcripts from YouTube, Instagram, TikTok, X, Vimeo and other video sites, summarize or analyze what was said (and shown on screen), or save an Obsidian-ready Markdown knowledge-base note…

    485 GitHub stars~3.1k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Volcengine Asr

    ysyecust/lecture-to-notes

    Transcribe local audio or video with Volcengine Doubao file ASR, including BigASR 1.0 Turbo direct upload and asynchronous 1.0 standard, 1.0 idle, or 2.0 standard jobs through TOS.

    273 GitHub stars~783 tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Openai Whisper

    huangruiteng/CS-Notes

    Local speech-to-text with the Whisper CLI (no API key). An agent skill from huangruiteng/CS-Notes.

    4k GitHub starsUsed in 18 repos~228 tokens
    AI & LLM EngineeringAuto-check passed
  • Ax Audio

    dosco/aithy

    This skill helps an LLM generate correct audio code with @ax-llm/ax.

    107 GitHub stars~2.5k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Stt Integration

    rapidaai/voice-ai

    Add or modify speech-to-text providers in assistant-api with transport-aware ingestion (WS/SDK/HTTP), transcript packet correctness, and UI/provider wiring.

    745 GitHub stars~778 tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed

More from Owl-Listener/inclusive-design-skills

All 55 skills in this repo
  • Assistive Technology Scenarios

    Owl-Listener/inclusive-design-skills

    Writes usage scenarios, use cases and storyboards that show real people using screen readers, switches, voice control and other assistive technology to finish tasks.

    104 GitHub stars~1.1k tokensUpdated 4 mo ago
    Auto-check passed
  • Error Prevention and Recovery

    Owl-Listener/inclusive-design-skills

    Designs forgiving forms and flows: prevent input errors, write messages that say what happened and what to do, and add undo, confirmation and recovery paths.

    104 GitHub stars~739 tokensUpdated 4 mo ago
    Auto-check passed
  • Accessible Heading Structure

    Owl-Listener/inclusive-design-skills

    Designs heading hierarchies for screen reader navigation and cognitive accessibility on pages, articles, dashboards and forms.

    104 GitHub stars~763 tokensUpdated 4 mo ago
    Auto-check passed
  • Ability Spectrum Mapping

    Owl-Listener/inclusive-design-skills

    Maps a feature across a range of vision, hearing, motor and cognitive ability to find where the design starts to fail, and to explain accessibility scope to stakeholders.

    104 GitHub stars~917 tokensUpdated 4 mo ago
    Auto-check passed
  • Accessibility Debt Tracking

    Owl-Listener/inclusive-design-skills

    Track and manage accessibility debt — known accessibility issues that have been deferred.

    104 GitHub stars~869 tokensUpdated 4 mo ago
    Auto-check passed
  • Accessibility Testing Strategy

    Owl-Listener/inclusive-design-skills

    Plan what to test, how to test, and who should test for accessibility.

    104 GitHub stars~1k tokensUpdated 4 mo ago
    Auto-check passed

Questions about Voice Interaction

What does Voice Interaction do?

Design voice interactions and speech interfaces that work for people with diverse speech patterns, accents, and communication styles. Voice Interaction is an agent skill from Owl-Listener/inclusive-design-skills. Design voice interactions and speech interfaces that work for people with diverse speech patterns, accents, and communication styles.

When should I use Voice Interaction?

Voice Interaction fits situations like: designing voice commands; voice assistants; any interface that accepts speech input; speech recognition.

How do I install Voice Interaction in Claude Code?

Run `npx skills add Owl-Listener/inclusive-design-skills --skill voice-interaction -a claude-code`. Or copy the skill folder (inclusive-interaction/skills/voice-interaction in Owl-Listener/inclusive-design-skills) into .claude/skills/voice-interaction in your project. Claude Code loads it when a task matches its description.

How do I install Voice Interaction in Codex?

Run `npx skills add Owl-Listener/inclusive-design-skills --skill voice-interaction -a codex`. Or copy the skill folder (inclusive-interaction/skills/voice-interaction in Owl-Listener/inclusive-design-skills) into .agents/skills/voice-interaction in your project. Codex loads it when a task matches its description.

Can I use Voice Interaction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Owl-Listener/inclusive-design-skills --skill voice-interaction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice-interaction, .gemini/skills/voice-interaction, .github/skills/voice-interaction and .opencode/skills/voice-interaction in your project.

What does Voice Interaction need to run?

SKILL.md names no scripts, command-line tools or credentials: Voice Interaction is instructions for the agent only.

Does Voice Interaction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Voice Interaction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voice Interaction use?

Voice Interaction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Voice Interaction use?

About 730 tokens (SKILL.md is roughly 2.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice Interaction?

Skills that share tags, products or a category with Voice Interaction: Yichen Asr (mcncarl/yichen-skills, 4.4k stars), Youtube Fetcher (JimmySadek/youtube-fetcher-to-markdown, 485 stars), Volcengine Asr (ysyecust/lecture-to-notes, 273 stars) and Openai Whisper (huangruiteng/CS-Notes, 4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice Interaction?

Owl-Listener (a GitHub user) maintains it in Owl-Listener/inclusive-design-skills, which has 104 GitHub stars. The repository holds 55 skills in this directory. The repository was last updated on June 9, 2026.

Source: Owl-Listener/inclusive-design-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.