Agent skill

Speech To Text

by NoizAI in NoizAI/skills

A skill your agent uses whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file.

No licenceAuto-check passedMedia & Creative

Install Speech To Text

skills CLI
$ npx skills add NoizAI/skills --skill speech-to-text -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NoizAI/skills speech-to-text --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NoizAI/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/speech-to-text .claude/skills/speech-to-text && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speech-to-text
GitHub stars
526
Token cost
~917 tokens
SKILL.md length
249 words
Files
2 (incl. scripts)
Skills in repo
8
Repo updated
First seen
Licence
None found

At a glance

A skill your agent uses whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file.

  • The user wants to transcribe audio to text
  • SKILL.md covers Triggers, Quick Start, Arguments and Output Format, plus 5 more sections
  • Runs Python scripts from its folder; calls python3 and pip; reaches noiz.ai; needs NOIZ_API_KEY
  • Convert speech to text

What it does

Speech To Text is an agent skill from NoizAI/skills. Use this skill whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file. Triggers include: any mention of 'transcribe', 'transcription', 'speech to text', 'STT', 'convert audio to text', 'what does this audio say', 'get transcript', 'subtitle generation', or requests to extract spoken words from a file. Also use when the user wants speaker identification from audio, timestamps for captions, or multilingual transcription.

Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/stt.py`).

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. The repository describes itself as: Allow your 🦞 bot to Shout, Speak, with "human" vibe.

When your agent uses it

  • The user wants to transcribe audio to text
  • Convert speech to text
  • Get a transcript from an audio
  • Include: any mention of transcribe

Example prompts

  • “transcribe”
  • “transcription”
  • “speech to text”
  • “/speech-to-text”

Requirements

  • Python 3
  • A credential in NOIZ_API_KEY
  • A credential in YOUR_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 779f30a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • noiz.ai

    Also links to:

    • developers.noiz.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NOIZ_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Speech To Text loads about 917 tokens when it runs. Until then it costs about 127 tokens; SKILL.md has 249 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~127
When it runs · the whole SKILL.md, loaded when a task matches
~917

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 249 words (~917 tokens).

“Transcribe any audio file to text. Supports multilingual auto-detection, timestamps, and speaker labels.”

— opening of SKILL.md by NoizAI
name
speech-to-text
permissions
network, filesystem

Read the full SKILL.md on GitHub

Files

SKILL.md and 1 other file (scripts) in skills/speech-to-text of NoizAI/skills.

  • SKILL.md
  • scripts/stt.py

Open the folder on GitHubat commit 779f30a

Compare with similar skills

Speech To Text next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speech To Text compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speech To Text this skillNoizAI/skills526—~917Automated safety check: PassNone
Voxtype Installpchalasani/claude-code-tools2k—~857Automated safety check: NotesMIT
Podcast Transcript FetcherVarnan-Tech/opendirectory674—~1.9kAutomated safety check: NotesMIT
Audio Transcriptionmitsuhiko/agent-stuff3.2k—~1kAutomated safety check: PassApache-2.0
Speech Recognitiondpearson2699/swift-ios-skills1.2k—~3.7kAutomated safety check: PassCustom licence
Speech Buildcnemri/google-genai-skills127—~430Automated safety check: PassMIT

Similar skills

  • Voxtype Install

    pchalasani/claude-code-tools

    Guide the user through installing, configuring, and launching voxtype — local on-device voice dictation (speech-to-text that types wherever the cursor is).

    2k GitHub stars~857 tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Podcast Transcript Fetcher

    Varnan-Tech/opendirectory

    A skill your agent uses when fetching, searching, or analyzing transcripts from Lenny's Podcast, Dwarkesh Podcast, Cheeky Pint, 20VC, or A16z Podcast.

    674 GitHub stars~1.9k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Audio Transcription

    mitsuhiko/agent-stuff

    Transcribe local audio/video and Apple Voice Memos quickly with cached MLX Whisper models, including bad/low-quality audio.

    3.2k GitHub stars~1k tokensUpdated 13 days ago
    Media & CreativeAuto-check passed
  • Speech Recognition

    dpearson2699/swift-ios-skills

    Transcribe speech to text using Apple's Speech framework. An agent skill from dpearson2699/swift-ios-skills.

    1.2k GitHub stars~3.7k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Speech Build

    cnemri/google-genai-skills

    Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models.

    127 GitHub stars~430 tokensUpdated 8 mo ago
    Media & CreativeAuto-check passed
  • Transcription Corrector

    cat-xierluo/legal-skills

    转录稿纠错与轻度优化。本技能应在用户需要按用户词典纠正 ASR 转录稿同音字与英文专有名称漂移时使用。不要用于:重写为课程章节、报告、总结,或完全空白的素材创作。

    720 GitHub stars~4.8k tokensUpdated yesterday
    Media & CreativeAuto-check passed

More from NoizAI/skills

All 8 skills in this repo
  • Music Video

    NoizAI/skills

    Build a music video from an existing song and lyrics, either as p5.js/p5.brush animation or by assembling local images and clips.

    526 GitHub stars~1.5k tokensUpdated 13 days ago
    Auto-check passed
  • A skill your agent uses whenever the user wants speech to sound more human, companion-like, or emotionally expressive.

    526 GitHub stars~1.8k tokensUpdated 13 days ago
    Auto-check passed
  • Chat With Anyone

    NoizAI/skills

    Chat with any real person or fictional character in their own voice by automatically finding their speech online, extracting a clean reference sample, and generating audio replies.

    526 GitHub stars~2.2k tokensUpdated 13 days ago
    Auto-check passed
  • Sound Fx

    NoizAI/skills

    A skill your agent uses whenever the user wants to generate sound effects, ambient audio, or short audio clips from a text description.

    526 GitHub stars~1.4k tokensUpdated 13 days ago
    Auto-check passed
  • Text To Music

    NoizAI/skills

    A skill your agent uses whenever the user wants to generate a song or music with vocals from a text description and/or lyrics, or cover an existing song in a new style.

    526 GitHub stars~2.2k tokensUpdated 13 days ago
    Auto-check passed
  • Tts

    NoizAI/skills

    A skill your agent uses whenever the user wants to convert text into speech, generate audio from text, or produce voiceovers.

    526 GitHub stars~1.9k tokensUpdated 13 days ago
    Auto-check passed

Questions about Speech To Text

What does Speech To Text do?

A skill your agent uses whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file. Speech To Text is an agent skill from NoizAI/skills. Use this skill whenever the user wants to transcribe audio to text, convert speech to text, or get a transcript from an audio or video file.

When should I use Speech To Text?

Speech To Text fits situations like: the user wants to transcribe audio to text; convert speech to text; get a transcript from an audio; include: any mention of transcribe.

How do I install Speech To Text in Claude Code?

Run `npx skills add NoizAI/skills --skill speech-to-text -a claude-code`. Or copy the skill folder (skills/speech-to-text in NoizAI/skills) into .claude/skills/speech-to-text in your project. Claude Code loads it when a task matches its description.

How do I install Speech To Text in Codex?

Run `npx skills add NoizAI/skills --skill speech-to-text -a codex`. Or copy the skill folder (skills/speech-to-text in NoizAI/skills) into .agents/skills/speech-to-text in your project. Codex loads it when a task matches its description.

Can I use Speech To Text in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NoizAI/skills --skill speech-to-text -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-to-text, .gemini/skills/speech-to-text, .github/skills/speech-to-text and .opencode/skills/speech-to-text in your project.

What does Speech To Text need to run?

Going by SKILL.md and its folder, Speech To Text needs Python for the scripts in its folder, the command-line tools its instructions call (python3 and pip) and credentials named NOIZ_API_KEY. Our summary lists: Python 3; A credential in NOIZ_API_KEY; A credential in YOUR_KEY.

Does Speech To Text access the network?

SKILL.md names 2 domains. In commands or code: noiz.ai; the agent is likely to contact it when it follows the instructions. As links in the text: developers.noiz.ai. This is read from the text; nothing was executed.

Is Speech To Text safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Speech To Text use?

No licence was found for Speech To Text or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Speech To Text use?

About 917 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Speech To Text?

Skills that share tags, products or a category with Speech To Text: Voxtype Install (pchalasani/claude-code-tools, 2k stars), Podcast Transcript Fetcher (Varnan-Tech/opendirectory, 674 stars), Audio Transcription (mitsuhiko/agent-stuff, 3.2k stars) and Speech Recognition (dpearson2699/swift-ios-skills, 1.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speech To Text?

NoizAI (a GitHub organization) maintains it in NoizAI/skills, which has 526 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 28, 2026.

Source: NoizAI/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.