Agent skill

Speech Recognition

by dpearson2699 in dpearson2699/swift-ios-skills

Transcribe speech to text using Apple's Speech framework. An agent skill from dpearson2699/swift-ios-skills.

Custom licenceAuto-check passedMedia & Creative

Install Speech Recognition

skills CLI
$ npx skills add dpearson2699/swift-ios-skills --skill speech-recognition -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dpearson2699/swift-ios-skills speech-recognition --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dpearson2699/swift-ios-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/speech-recognition .claude/skills/speech-recognition && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
speech-recognition
GitHub stars
1.2k
Token cost
~3.7k tokens
SKILL.md length
788 words
Files
3 (incl. references)
Skills in repo
86
Repo updated
First seen
Licence
Custom licence

At a glance

Transcribe speech to text using Apple's Speech framework. An agent skill from dpearson2699/swift-ios-skills.

  • Works in 7 steps: Choose the module → Check support before creating the session → Pick a documented preset → …
  • Implementing live microphone transcription with AVAudioEngine
  • SKILL.md covers Contents, SpeechAnalyzer Strategy (iOS…, SFSpeechRecognizer Setup and Authorization, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Speech Recognition is an agent skill from dpearson2699/swift-ios-skills. Transcribe speech to text using Apple's Speech framework. Use when implementing live microphone transcription with AVAudioEngine, recognizing recorded audio files, handling speech and microphone authorization, choosing on-device vs server-backed SFSpeechRecognizer behavior, or adopting SpeechAnalyzer, SpeechTranscriber, DictationTranscriber, AssetInventory, and async result streams on iOS 26+.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `evals/evals.json` and `references/speechanalyzer-patterns.md`).

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with iOS. The repository describes itself as: Agent Skills for iOS 26+, Swift 6.3, SwiftUI, and modern Apple frameworks.

When your agent uses it

  • Implementing live microphone transcription with AVAudioEngine
  • Recognizing recorded audio files
  • Handling speech and microphone authorization
  • Choosing on-device vs server-backed SFSpeechRecognizer behavior

Example prompts

  • “/speech-recognition”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Choose the module
  2. Check support before creating the session
  3. Pick a documented preset
  4. Install required assets with AssetInventory.assetInstallationRequest.
  5. Convert live audio buffers to
  6. Consume module results from their AsyncSequence in a separate task.
  7. Finish explicitly with finalizeAndFinish(through:),

What it can do on your machine

Read from SKILL.md and the folder at commit 8d90fd1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are swift).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • sosumi.ai

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Speech Recognition loads about 3.7k tokens when it runs, and up to ~5.2k if it reads all its reference files. Until then it costs about 104 tokens; SKILL.md has 788 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~104
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~5.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 788 words (~3,652 tokens).

“Transcribe live and pre-recorded audio to text using Apple's Speech framework. Covers SpeechAnalyzer / SpeechTranscriber (iOS 26+) and SFSpeechRecognizer (iOS 10+) fallback guidance.”

— opening of SKILL.md by dpearson2699, Custom licence
name
speech-recognition

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files (references) in skills/speech-recognition of dpearson2699/swift-ios-skills.

  • SKILL.md
  • evals/evals.json
  • references/speechanalyzer-patterns.md

Open the folder on GitHubat commit 8d90fd1

Compare with similar skills

Speech Recognition next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Speech Recognition compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Speech Recognition this skilldpearson2699/swift-ios-skills1.2k—~3.7kAutomated safety check: PassCustom licence
WhisperAlexAI-MCP/hermes-CCC135—~1.9kAutomated safety check: PassMIT
Muapi AI Clippingmajiayu000/claude-skill-registry6661 repos~1.7kAutomated safety check: PassMIT
Openai Whisper APICoWork-OS/CoWork-OS473—~411Automated safety check: PassMIT
9Router Speech-to-Textdecolua/9router30k—~745Automated safety check: PassMIT
Openai Whisper APIopenclaw/openclaw392k2 repos~518Automated safety check: PassMIT

Similar skills

  • Whisper

    AlexAI-MCP/hermes-CCC

    OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.

    135 GitHub stars~1.9k tokensUpdated 6 mo ago
    Media & CreativeAuto-check passed
  • Muapi AI Clipping

    majiayu000/claude-skill-registry

    Turn a long video into N viral-ready short clips with a single managed API call.

    666 GitHub starsUsed in 1 repo~1.7k tokens
    Media & CreativeAuto-check passed
  • Openai Whisper API

    CoWork-OS/CoWork-OS

    Transcribe audio via OpenAI Whisper, Atlas Cloud, or MuAPI speech-to-text APIs.

    473 GitHub stars~411 tokensUpdated today
    Media & CreativeAuto-check passed
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    30k GitHub stars~745 tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Openai Whisper API

    openclaw/openclaw

    OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

    392k GitHub starsUsed in 2 repos~518 tokens
    Media & CreativeAuto-check passed
  • Life Recorder Setup

    browser-use/life-recorder

    Build, install, pair, and operate the Life Recorder iPhone-to-Mac local transcription system.

    321 GitHub stars~599 tokensUpdated 27 days ago
    Media & CreativeAuto-check passed

More from dpearson2699/swift-ios-skills

All 86 skills in this repo
  • iOS Memgraph Analysis

    dpearson2699/swift-ios-skills

    A skill your agent uses when capturing or analyzing an iOS .memgraph, especially when the task mentions a memory leak, heap growth, persistent memory increase, ownership path, or matched-capture…

    1.2k GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Natural Language

    dpearson2699/swift-ios-skills

    Tokenize, tag, and analyze natural language text using Apple's NaturalLanguage framework and translate between languages with the Translation framework.

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Auto-check passed
  • iOS Ettrace Performance

    dpearson2699/swift-ios-skills

    A skill your agent uses when capturing or analyzing ETTrace profiles for a focused iOS launch or runtime flow, including exact-build dSYM UUID matching, Simulator or device capture, processed…

    1.2k GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Accessorysetupkit

    dpearson2699/swift-ios-skills

    Discover and configure Bluetooth and Wi-Fi accessories using AccessorySetupKit.

    1.2k GitHub stars~3.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Activitykit

    dpearson2699/swift-ios-skills

    Implement, review, or improve Live Activities and Dynamic Island experiences in iOS apps using ActivityKit.

    1.2k GitHub stars~4.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Adattributionkit

    dpearson2699/swift-ios-skills

    Measure ad effectiveness with privacy-preserving attribution using AdAttributionKit.

    1.2k GitHub stars~3.5k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Speech Recognition

What does Speech Recognition do?

Transcribe speech to text using Apple's Speech framework. An agent skill from dpearson2699/swift-ios-skills. Speech Recognition is an agent skill from dpearson2699/swift-ios-skills. Transcribe speech to text using Apple's Speech framework.

When should I use Speech Recognition?

Speech Recognition fits situations like: implementing live microphone transcription with AVAudioEngine; recognizing recorded audio files; handling speech and microphone authorization; choosing on-device vs server-backed SFSpeechRecognizer behavior.

How do I install Speech Recognition in Claude Code?

Run `npx skills add dpearson2699/swift-ios-skills --skill speech-recognition -a claude-code`. Or copy the skill folder (skills/speech-recognition in dpearson2699/swift-ios-skills) into .claude/skills/speech-recognition in your project. Claude Code loads it when a task matches its description.

How do I install Speech Recognition in Codex?

Run `npx skills add dpearson2699/swift-ios-skills --skill speech-recognition -a codex`. Or copy the skill folder (skills/speech-recognition in dpearson2699/swift-ios-skills) into .agents/skills/speech-recognition in your project. Codex loads it when a task matches its description.

Can I use Speech Recognition in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dpearson2699/swift-ios-skills --skill speech-recognition -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/speech-recognition, .gemini/skills/speech-recognition, .github/skills/speech-recognition and .opencode/skills/speech-recognition in your project.

What does Speech Recognition need to run?

SKILL.md names no scripts, command-line tools or credentials: Speech Recognition is instructions for the agent only.

Does Speech Recognition access the network?

SKILL.md names 1 domain. As links in the text: sosumi.ai. This is read from the text; nothing was executed.

Is Speech Recognition safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Speech Recognition use?

Speech Recognition has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Speech Recognition use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.5k tokens, read only when the agent opens those files.

What are the alternatives to Speech Recognition?

Skills that share tags, products or a category with Speech Recognition: Whisper (AlexAI-MCP/hermes-CCC, 135 stars), Muapi AI Clipping (majiayu000/claude-skill-registry, 666 stars), Openai Whisper API (CoWork-OS/CoWork-OS, 473 stars) and 9Router Speech-to-Text (decolua/9router, 30k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Speech Recognition?

dpearson2699 (a GitHub user) maintains it in dpearson2699/swift-ios-skills, which has 1,175 GitHub stars. The repository holds 86 skills in this directory. The repository was last updated on July 31, 2026.

Source: dpearson2699/swift-ios-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.