Agent skill

Openai Whisper API

by openclaw in openclaw/openclaw

OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

MITAuto-check passedMedia & Creative

Install Openai Whisper API

skills CLI
$ npx skills add openclaw/openclaw --skill openai-whisper-api -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install openclaw/openclaw openai-whisper-api --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/openclaw/openclaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/openai-whisper-api .claude/skills/openai-whisper-api && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
openai-whisper-api
GitHub stars
392k
Used in
1 other repo
Token cost
~518 tokens
SKILL.md length
75 words
Files
2 (incl. scripts)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

  • Tasks that involve Transcription
  • SKILL.md covers Quick start, Useful flags and API key
  • Runs Shell scripts from its folder; needs OPENAI_API_KEY
  • Tasks that involve Speech recognition and synthesis

What it does

Openai Whisper API is an agent skill from openclaw/openclaw. OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

Its SKILL.md is about 520 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/transcribe.sh`).

It sits in Media & Creative, covering Transcription and Speech recognition and synthesis. It works with OpenAI and Whisper. The repository describes itself as: The AI that really does things. Any OS. Any Platform. The lobster way. 🦞. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Speech recognition and synthesis

Example prompts

  • “/openai-whisper-api”

Requirements

  • A Bash shell
  • A credential in OPENAI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit f8e594e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Openai Whisper API loads about 518 tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 75 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~518

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from openclaw/openclaw at commit f8e594e, republished under its MIT licence (© openclaw). 75 words, ~518 tokens.

Download SKILL.mdSave it as .claude/skills/openai-whisper-api/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
openai-whisper-api
description
OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.
homepage
https://platform.openai.com/docs/guides/speech-to-text

OpenAI transcriptions API

Transcribe audio through /v1/audio/transcriptions. Set OPENAI_BASE_URL for an OpenAI-compatible proxy or local gateway.

Quick start

bash
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a

Defaults:

  • Model: gpt-4o-transcribe
  • Output: <input>.txt

Useful flags

bash
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model gpt-4o-transcribe --out /tmp/transcript.txt
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model gpt-4o-mini-transcribe
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model gpt-4o-transcribe-diarize --json
{baseDir}/scripts/transcribe.sh /path/to/audio.ogg --model whisper-1
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --language en
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --prompt "Speaker names: Peter, Daniel"
{baseDir}/scripts/transcribe.sh /path/to/audio.m4a --json --out /tmp/transcript.json

Notes:

  • Supported upload formats include mp3, mp4, mpeg, mpga, m4a, wav, webm.
  • 25 MB upload limit on the hosted API.
  • Use diarize for speaker labels; script sends chunking_strategy=auto and rejects --prompt.

API key

Set OPENAI_API_KEY, or configure it in the active OpenClaw config file ($OPENCLAW_CONFIG_PATH, default ~/.openclaw/openclaw.json). Optionally set OPENAI_BASE_URL:

json5
{
  skills: {
    "openai-whisper-api": {
      apiKey: "OPENAI_KEY_HERE",
    },
  },
}

© openclaw, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in skills/openai-whisper-api of openclaw/openclaw.

  • SKILL.md
  • scripts/transcribe.sh

Open the folder on GitHubat commit f8e594e

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in openclaw/openclaw, which our catalogue first saw on October 8, 2026.

Compare with similar skills

Openai Whisper API next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Openai Whisper API compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Openai Whisper API this skillopenclaw/openclaw392k1 repos~518Automated safety check: PassMIT
9Router Speech-to-Textdecolua/9router30k—~914Automated safety check: PassMIT
Openai Whispercoco-research/coco503—~964Automated safety check: PassCustom licence
Openai Whisper APItrpc-group/trpc-agent-go1.9k12 repos~288Automated safety check: PassApache-2.0
Keirouter Sttmydisha/keirouter147—~680Automated safety check: PassMIT
Whisper Speech RecognitionOrchestra-Research/AI-Research-SKILLs13k7 repos~1.9kAutomated safety check: NotesMIT

Similar skills

  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    30k GitHub stars~914 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Openai Whisper

    coco-research/coco

    Speech-to-text transcription via OpenAI Whisper. An agent skill from coco-research/coco.

    503 GitHub stars~964 tokensUpdated today
    Media & CreativeAuto-check passed
  • Openai Whisper API

    trpc-group/trpc-agent-go

    Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

    1.9k GitHub starsUsed in 12 repos~288 tokens
    AI & LLM EngineeringAuto-check passed
  • Keirouter Stt

    mydisha/keirouter

    Speech-to-text via KeiRouter /v1/audio/transcriptions using OpenAI Whisper / Groq / Gemini / Deepgram / AssemblyAI models.

    147 GitHub stars~680 tokensUpdated 28 days ago
    Media & CreativeAuto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    AI & LLM EngineeringAuto-check: notes
  • Local AI Use

    amd/skills

    Makes this agent generate images, transcribe audio, and synthesize speech on the user's own machine through a local Lemonade Server instead of a paid cloud API.

    406 GitHub stars~5k tokensUpdated today
    Media & CreativeAuto-check: notes

More from openclaw/openclaw

All 96 skills in this repo
  • Model Usage

    openclaw/openclaw

    Summarize CodexBar local cost logs by model for Codex or Claude, including current or full breakdowns.

    392k GitHub starsUsed in 1 repo~637 tokens
    Auto-check passed
  • Openclaw Live Updater

    openclaw/openclaw

    Maintain the canonical live OpenClaw main checkout, macOS LaunchAgent-managed Gateway, local macOS app, exact-head main CI, and recurring full release validation.

    392k GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Feishu Doc

    openclaw/openclaw

    Feishu document read/write workflows. An agent skill from openclaw/openclaw.

    392k GitHub stars~516 tokensUpdated today
    Auto-check passed
  • Tmux

    openclaw/openclaw

    Control tmux sessions/panes for interactive CLIs: list, capture output, send keys, paste text, monitor prompts.

    392k GitHub starsUsed in 1 repo~640 tokens
    Auto-check passed
  • Openclaw PR Maintainer

    openclaw/openclaw

    Review, triage, repair, or land OpenClaw issues and pull requests with current-source evidence and the native maintainer workflow.

    392k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Browser Automation

    openclaw/openclaw

    A skill your agent uses when controlling web pages with the OpenClaw browser tool, especially multi-step flows, login checks, tab management, or recovery from stale refs/timeouts.

    392k GitHub stars~2.9k tokensUpdated today
    Auto-check passed

Works with

Questions about Openai Whisper API

What does Openai Whisper API do?

OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1. Openai Whisper API is an agent skill from openclaw/openclaw. OpenAI Audio Transcriptions API via curl; gpt-4o-transcribe, mini, diarize, or whisper-1.

When should I use Openai Whisper API?

Openai Whisper API fits situations like: tasks that involve Transcription; tasks that involve Speech recognition and synthesis.

How do I install Openai Whisper API in Claude Code?

Run `npx skills add openclaw/openclaw --skill openai-whisper-api -a claude-code`. Or copy the skill folder (skills/openai-whisper-api in openclaw/openclaw) into .claude/skills/openai-whisper-api in your project. Claude Code loads it when a task matches its description.

How do I install Openai Whisper API in Codex?

Run `npx skills add openclaw/openclaw --skill openai-whisper-api -a codex`. Or copy the skill folder (skills/openai-whisper-api in openclaw/openclaw) into .agents/skills/openai-whisper-api in your project. Codex loads it when a task matches its description.

Can I use Openai Whisper API in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openclaw/openclaw --skill openai-whisper-api -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/openai-whisper-api, .gemini/skills/openai-whisper-api, .github/skills/openai-whisper-api and .opencode/skills/openai-whisper-api in your project.

What does Openai Whisper API need to run?

Going by SKILL.md and its folder, Openai Whisper API needs a shell for the scripts in its folder and credentials named OPENAI_API_KEY. Our summary lists: A Bash shell; A credential in OPENAI_API_KEY.

Does Openai Whisper API access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Openai Whisper API safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Openai Whisper API use?

Openai Whisper API is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Openai Whisper API use?

About 518 tokens (SKILL.md is roughly 2.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Openai Whisper API?

Skills that share tags, products or a category with Openai Whisper API: 9Router Speech-to-Text (decolua/9router, 30k stars), Openai Whisper (coco-research/coco, 503 stars), Openai Whisper API (trpc-group/trpc-agent-go, 1.9k stars) and Keirouter Stt (mydisha/keirouter, 147 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Openai Whisper API?

openclaw (a GitHub organization) maintains it in openclaw/openclaw, which has 391,504 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on October 9, 2026.

Source: openclaw/openclaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.