Agent skill

Voice Transcription

by BuilderIO in BuilderIO/agent-native

Framework-wide voice dictation in the agent sidebar composer.

No licenceAuto-check passedMedia & Creative

Install Voice Transcription

skills CLI
$ npx skills add BuilderIO/agent-native --skill voice-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BuilderIO/agent-native voice-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BuilderIO/agent-native.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/voice-transcription .claude/skills/voice-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
voice-transcription
GitHub stars
7.1k
Token cost
~4.2k tokens
SKILL.md length
1,837 words
Files
1
Skills in repo
108
Repo updated
First seen
Licence
None found

At a glance

Framework-wide voice dictation in the agent sidebar composer.

  • Works in 5 steps: If builder-gemini and… → If builder and… → If gemini → resolves the Gemini key… → …
  • Changing composer microphone UX
  • SKILL.md covers UX rules, Realtime speech mode, Source And Cleanup and Where the pieces live, plus 3 more sections
  • Needs OPENAI_API_KEY and GOOGLE_GENERATIVE_AI_API_KEY

What it does

Voice Transcription is an agent skill from BuilderIO/agent-native. Framework-wide voice dictation in the agent sidebar composer. Use when changing composer microphone UX, the transcribe-voice route, or the Voice Transcription settings section. Covers transcription-source routing, cleanup routing, Google realtime gating, and the voice transcription application-state keys.

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription. The repository describes itself as: A framework for building agentic apps.

When your agent uses it

  • Changing composer microphone UX
  • The transcribe-voice route
  • The Voice Transcription settings section

Example prompts

  • “/voice-transcription”

Requirements

  • A credential in OPENAI_API_KEY
  • A credential in GOOGLE_GENERATIVE_AI_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. If builder-gemini and resolveHasBuilderPrivateKey() → calls transcribeWithBuilder({ model: "gemini-3-1-flash-lite" }) via Builder proxy…
  2. If builder and resolveHasBuilderPrivateKey() → legacy alias; prefer builder-gemini.
  3. If gemini → resolves the Gemini key under either name (resolveGeminiApiKey) and calls the direct Google Gemini path.
  4. If groq → resolves GROQ_API_KEY and calls Groq's Whisper-compatible endpoint.
  5. If openai → resolves OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 1c2de07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • cloud.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • GOOGLE_GENERATIVE_AI_API_KEY
    • GEMINI_API_KEY
    • GOOGLE_APPLICATION_CREDENTIALS
    • GROQ_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Voice Transcription loads about 4.2k tokens when it runs. Until then it costs about 82 tokens; SKILL.md has 1,837 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~82
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,837 words (~4,150 tokens).

“The microphone inside the sidebar composer offers two distinct paths: editable dictation and an opt-in realtime speech-to-speech agent session. Users configure dictation separately from AI cleanup. The source picker (Mac Native, Google Realtime, Batch) is the Voice transcription row on…”

— opening of SKILL.md by BuilderIO
name
voice-transcription
scope
dev
metadata.internal
true

Read the full SKILL.md on GitHub

Files

Just SKILL.md in .agents/skills/voice-transcription of BuilderIO/agent-native.

Open the folder on GitHubat commit 1c2de07

Compare with similar skills

Voice Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Voice Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Voice Transcription this skillBuilderIO/agent-native7.1k—~4.2kAutomated safety check: PassNone
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.6k—~1.8kAutomated safety check: PassMIT
Edu Chem Videowy51ai/edulab1.4k—~2.1kAutomated safety check: NotesApache-2.0
Transcription Memory ReconstructionNxcoreAI/EverRoom3k—~714Automated safety check: PassCustom licence
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.6k GitHub stars~1.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Edu Chem Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a chemistry problem (化学题: 氧化还原配平 双线桥 电子守恒, 物质的量计算, 化学平衡 三段式 平衡常数 转化率 反应速率, 离子反应, 电化学, 溶液 滴定…

    1.4k GitHub stars~2.1k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Reconstruct a complete, searchable memory from an untrusted meeting or conversation transcript.

    3k GitHub stars~714 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated today
    Media & CreativeAuto-check: notes
  • Transcribe

    JetBrains/skills

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints.

    366 GitHub starsUsed in 4 repos~776 tokens
    Media & CreativeAuto-check passed

More from BuilderIO/agent-native

All 108 skills in this repo
  • Bug Trace

    BuilderIO/agent-native

    Find the root cause of a reported bug: take in the report (GitHub issue, Jira ticket, file, or pasted text), trace the failing path hop by hop with each hop's runtime precondition, find the recent…

    7.1k GitHub stars~555 tokensUpdated today
    Auto-check passed
  • Identify Fragile Systems

    BuilderIO/agent-native

    Nightly refactor review: find the systems the last day's commits hit hardest, use three weeks of history to tell fragile from fast-moving, write a plan per systemic fix, and file deduplicated Jira…

    7.1k GitHub stars~957 tokensUpdated today
    Auto-check passed
  • Jira Refactor Findings

    BuilderIO/agent-native

    File and track refactor findings in Jira without duplicates: match a run against existing refactor-findings tickets, create tickets from plan files (Pod, label, attached plan, run link), record…

    7.1k GitHub stars~779 tokensUpdated today
    Auto-check passed
  • System History

    BuilderIO/agent-native

    Blobless-safe git and PR history for this repo. An agent skill from BuilderIO/agent-native.

    7.1k GitHub stars~727 tokensUpdated today
    Auto-check passed
  • Actions

    BuilderIO/agent-native

    How to create and run agent actions. An agent skill from BuilderIO/agent-native.

    7.1k GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Client Methods

    BuilderIO/agent-native

    Client method surface rules. An agent skill from BuilderIO/agent-native.

    7.1k GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about Voice Transcription

What does Voice Transcription do?

Framework-wide voice dictation in the agent sidebar composer. Voice Transcription is an agent skill from BuilderIO/agent-native. Framework-wide voice dictation in the agent sidebar composer.

When should I use Voice Transcription?

Voice Transcription fits situations like: changing composer microphone UX; the transcribe-voice route; the Voice Transcription settings section.

How do I install Voice Transcription in Claude Code?

Run `npx skills add BuilderIO/agent-native --skill voice-transcription -a claude-code`. Or copy the skill folder (.agents/skills/voice-transcription in BuilderIO/agent-native) into .claude/skills/voice-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Voice Transcription in Codex?

Run `npx skills add BuilderIO/agent-native --skill voice-transcription -a codex`. Or copy the skill folder (.agents/skills/voice-transcription in BuilderIO/agent-native) into .agents/skills/voice-transcription in your project. Codex loads it when a task matches its description.

Can I use Voice Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BuilderIO/agent-native --skill voice-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/voice-transcription, .gemini/skills/voice-transcription, .github/skills/voice-transcription and .opencode/skills/voice-transcription in your project.

What does Voice Transcription need to run?

Going by SKILL.md and its folder, Voice Transcription needs credentials named OPENAI_API_KEY, GOOGLE_GENERATIVE_AI_API_KEY, GEMINI_API_KEY and GOOGLE_APPLICATION_CREDENTIALS. Our summary lists: A credential in OPENAI_API_KEY; A credential in GOOGLE_GENERATIVE_AI_API_KEY.

Does Voice Transcription access the network?

SKILL.md names 1 domain. As links in the text: cloud.google.com. This is read from the text; nothing was executed.

Is Voice Transcription safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Voice Transcription use?

No licence was found for Voice Transcription or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Voice Transcription use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Voice Transcription?

Skills that share tags, products or a category with Voice Transcription: HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.6k stars), Edu Chem Video (wy51ai/edulab, 1.4k stars) and Transcription Memory Reconstruction (NxcoreAI/EverRoom, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Voice Transcription?

BuilderIO (a GitHub organization) maintains it in BuilderIO/agent-native, which has 7,113 GitHub stars. The repository holds 108 skills in this directory. The repository was last updated on October 10, 2026.

Source: BuilderIO/agent-native on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.