Agent skill

Yao Audio Transcription

by YaoApp in YaoApp/yao

Transcribes audio files such as mp3, m4a, wav and webm to text with the tai audio_transcribe tool, and lists the available speech-to-text providers.

Custom licenceAuto-check passedMedia & Creative

Install Yao Audio Transcription

skills CLI
$ npx skills add YaoApp/yao --skill yao-audio -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install YaoApp/yao yao-audio --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/YaoApp/yao.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tools/skills/yao-audio .claude/skills/yao-audio && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
yao-audio
GitHub stars
8.1k
Token cost
~416 tokens
SKILL.md length
130 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Custom licence

At a glance

Transcribes audio files such as mp3, m4a, wav and webm to text with the tai audio_transcribe tool, and lists the available speech-to-text providers.

  • Transcribing a meeting recording to text
  • SKILL.md covers audio_transcribe, audio_providers and Constraints
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Converting a voice memo or speech file into written text

What it does

Two tools are documented. The audio_transcribe tool takes an audio file path, an optional ISO 639-1 language code that is auto-detected when omitted, and an optional provider connector ID that falls back to the default speech-to-text provider. Supported formats are mp3, m4a, wav, webm, mp4, mpeg and mpga, and it is called through the tai tool command line.

The audio_providers tool lists the available speech-to-text providers with their models and connector IDs, optionally filtered by capability, so a provider can be passed on to audio_transcribe. The skill insists on using only the listed parameters, since others are ignored or cause errors.

When your agent uses it

  • Transcribing a meeting recording to text
  • Converting a voice memo or speech file into written text
  • Choosing which speech-to-text provider to use

Example prompts

  • “Transcribe /recordings/meeting.m4a to text.”
  • “Which speech-to-text providers are available?”
  • “Transcribe this Japanese voice memo and set the language to ja.”

Requirements

  • The tai command line tool
  • A configured speech-to-text provider

What it can do on your machine

Read from SKILL.md and the folder at commit eea1e7e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Yao Audio Transcription loads about 416 tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 130 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~416

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 130 words (~416 tokens).

“Use these tools to transcribe audio files to text using speech-to-text models.”

— opening of SKILL.md by YaoApp, Custom licence
name
yao-audio

Read the full SKILL.md on GitHub

Files

Just SKILL.md in tools/skills/yao-audio of YaoApp/yao.

Open the folder on GitHubat commit eea1e7e

Compare with similar skills

Yao Audio Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Yao Audio Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Yao Audio Transcription this skillYaoApp/yao8.1k—~416Automated safety check: PassCustom licence
Watch Video Q&Abradautomates/claude-video18k—~4.3kAutomated safety check: NotesMIT
Claude Real VideoHUANGCHIHHUNGLeo/claude-real-video2.2k—~639Automated safety check: PassMIT
9Router Speech-to-Textdecolua/9router30k—~914Automated safety check: PassMIT
Claude Real Video For AgentsHUANGCHIHHUNGLeo/claude-real-video2.2k—~2kAutomated safety check: NotesMIT
Karaoke CaptionsAI-Builder-Club/skills1.3k—~850Automated safety check: PassNone

Similar skills

  • Watch Video Q&A

    bradautomates/claude-video

    Lets the agent answer questions about a video from a URL or local file by downloading it, extracting frames and a transcript, or by sending it to Gemini's video model.

    18k GitHub stars~4.3k tokensUpdated 14 days ago
    Media & CreativeAuto-check: notes
  • Claude Real Video

    HUANGCHIHHUNGLeo/claude-real-video

    Watch a video for the user. An agent skill from HUANGCHIHHUNGLeo/claude-real-video.

    2.2k GitHub stars~639 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    30k GitHub stars~914 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Claude Real Video For Agents

    HUANGCHIHHUNGLeo/claude-real-video

    Install and use crv (claude-real-video) — a tool that lets any AI agent watch videos by extracting scene-aware keyframes, deduplicating them, and transcribing audio.

    2.2k GitHub stars~2k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Karaoke Captions

    AI-Builder-Club/skills

    Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass.

    1.3k GitHub stars~850 tokensUpdated 23 days ago
    Media & CreativeAuto-check passed
  • Speech To Text

    tadaspetra/loop

    Transcribe audio to text using ElevenLabs Scribe v2. An agent skill from tadaspetra/loop.

    296 GitHub starsUsed in 2 repos~2k tokens
    Media & CreativeAuto-check passed

More from YaoApp/yao

All 14 skills in this repo
  • Lists, downloads, references and deploys agents on a Yao host through five bash-invoked tools, keeping edits to the smith namespace while allowing read-only study of any agent.

    8.1k GitHub stars~873 tokensUpdated yesterday
    Auto-check passed
  • Starts, lists, checks, restarts and stops long-running background services through Yao daemon tools, reporting the verified port and URL of each.

    8.1k GitHub stars~1.1k tokensUpdated yesterday
    Auto-check passed
  • Yao Image Tools

    YaoApp/yao

    Lets an agent read and describe images through a vision model and generate or edit images from text prompts, using the tai command-line tool.

    8.1k GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Starts, lists, monitors, waits on and stops long-running background commands through a dedicated set of job-management tools.

    8.1k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Lists workspaces and reads or writes their files on a remote node through five tai tool commands, for cross-node file work the local filesystem cannot do.

    8.1k GitHub stars~939 tokensUpdated yesterday
    Auto-check passed
  • Manages Git identity, HTTPS access tokens and SSH keys at the workspace level through three `tai tool` commands, each with a get, set, list, import or delete action.

    8.1k GitHub stars~712 tokensUpdated yesterday
    Auto-check passed

Questions about Yao Audio Transcription

What does Yao Audio Transcription do?

Transcribes audio files such as mp3, m4a, wav and webm to text with the tai audio_transcribe tool, and lists the available speech-to-text providers. Two tools are documented. The audio_transcribe tool takes an audio file path, an optional ISO 639-1 language code that is auto-detected when omitted, and an optional provider connector ID that falls back to the default speech-to-text provider.

When should I use Yao Audio Transcription?

Yao Audio Transcription fits situations like: transcribing a meeting recording to text; converting a voice memo or speech file into written text; choosing which speech-to-text provider to use.

How do I install Yao Audio Transcription in Claude Code?

Run `npx skills add YaoApp/yao --skill yao-audio -a claude-code`. Or copy the skill folder (tools/skills/yao-audio in YaoApp/yao) into .claude/skills/yao-audio in your project. Claude Code loads it when a task matches its description.

How do I install Yao Audio Transcription in Codex?

Run `npx skills add YaoApp/yao --skill yao-audio -a codex`. Or copy the skill folder (tools/skills/yao-audio in YaoApp/yao) into .agents/skills/yao-audio in your project. Codex loads it when a task matches its description.

Can I use Yao Audio Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add YaoApp/yao --skill yao-audio -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/yao-audio, .gemini/skills/yao-audio, .github/skills/yao-audio and .opencode/skills/yao-audio in your project.

What does Yao Audio Transcription need to run?

SKILL.md names no scripts, command-line tools or credentials: Yao Audio Transcription is instructions for the agent only. Our summary lists: The tai command line tool; A configured speech-to-text provider.

Does Yao Audio Transcription access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Yao Audio Transcription safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Yao Audio Transcription use?

Yao Audio Transcription has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Yao Audio Transcription use?

About 416 tokens (SKILL.md is roughly 1.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Yao Audio Transcription?

Skills that share tags, products or a category with Yao Audio Transcription: Watch Video Q&A (bradautomates/claude-video, 18k stars), Claude Real Video (HUANGCHIHHUNGLeo/claude-real-video, 2.2k stars), 9Router Speech-to-Text (decolua/9router, 30k stars) and Claude Real Video For Agents (HUANGCHIHHUNGLeo/claude-real-video, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Yao Audio Transcription?

YaoApp (a GitHub organization) maintains it in YaoApp/yao, which has 8,103 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on October 8, 2026.

Source: YaoApp/yao on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.