Agent skill

Audio To Subtitles

by bozhouDev in bozhouDev/video-skills-toolkit

Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API.

MITAuto-check: notesMedia & Creative

Install Audio To Subtitles

skills CLI
$ npx skills add bozhouDev/video-skills-toolkit --skill audio-to-subtitles -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bozhouDev/video-skills-toolkit audio-to-subtitles --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bozhouDev/video-skills-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/audio-to-subtitles .claude/skills/audio-to-subtitles && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
audio-to-subtitles
GitHub stars
150
Token cost
~1.9k tokens
SKILL.md length
657 words
Files
3 (incl. scripts)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API.

  • Works in 4 steps: If the input is a local file, upload it… → Submit the public audio_url or video_url… → Poll the async task until it is… → …
  • The user asks to turn audio/video into subtitles
  • SKILL.md covers Workflow, CLI, Options and Manifest, plus 2 more sections
  • Runs TypeScript scripts from its folder; calls npx; reaches api.minimax.io; needs MINIMAX_API_KEY and MINIMAX_TTS_API_KEY

What it does

Audio To Subtitles is an agent skill from bozhouDev/video-skills-toolkit. Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API. Also supports MiniMax text-to-speech generation from text or text files, with optional subtitle generation from the produced audio. Supports batch processing, SRT/VTT/JSON output, language selection, speaker labels, configurable MiniMax voice ID and model. Use when the user asks to turn audio/video into subtitles, generate ASR subtitles, create…

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including scripts (for example `scripts/main.ts` and `scripts/r2.ts`).

It sits in Media & Creative, covering Transcription and Text to speech and voice. It works with MiniMax, Cloudflare R2 and HeyGen. The repository describes itself as: Video skills toolkit for Remotion talking-head, sketch story, and audio-to-subtitles workflows. The licence is MIT.

When your agent uses it

  • The user asks to turn audio/video into subtitles
  • Generate ASR subtitles
  • Batch transcription
  • Generate TTS audio with optional subtitles

Example prompts

  • “/audio-to-subtitles”

Requirements

  • Node.js
  • A credential in MINIMAX_API_KEY
  • A credential in MINIMAX_TTS_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. If the input is a local file, upload it to Cloudflare R2 first.
  2. Submit the public audio_url or video_url to AI MediaKit ASR subtitles.
  3. Poll the async task until it is completed or failed.
  4. Save .json, .srt, and .vtt outputs.

What it can do on your machine

Read from SKILL.md and the folder at commit 4766a16. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • npx

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.minimax.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • MINIMAX_API_KEY
    • MINIMAX_TTS_API_KEY
    • MEDIAKIT_API_KEY
    • AI_MEDIAKIT_API_KEY
    • VOLCENGINE_MEDIAKIT_API_KEY
    • R2_ACCESS_KEY_ID
    • R2_SECRET_ACCESS_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Audio To Subtitles loads about 1.9k tokens when it runs. Until then it costs about 153 tokens; SKILL.md has 657 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~153
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:113
    onment files in this order. The nearest `.env.r2` is loaded last and overrides global R2 defaults for this vault.
  • NoteMentions a .env fileSKILL.md:115
    1. `~/.skills/.env`
  • NoteMentions a .env fileSKILL.md:116
    2. `~/.baoyu-skills/.env`
  • NoteMentions a .env fileSKILL.md:117
    3. nearest `.env.r2` from current directory upward
  • NoteMentions a .env fileSKILL.md:118
    4. nearest `.skills/.env` from current directory upward
  • NoteMentions a .env fileSKILL.md:119
    5. nearest `.baoyu-skills/.env` from current directory upward
  • NoteMentions a .env fileSKILL.md:157
    - Do not print API keys or commit real `.env.r2` values.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bozhouDev/video-skills-toolkit at commit 4766a16, republished under its MIT licence (© bozhouDev). 657 words, ~1,892 tokens.

Download SKILL.mdSave it as .claude/skills/audio-to-subtitles/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
audio-to-subtitles
description
Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API. Also supports MiniMax text-to-speech generation from text or text files, with optional subtitle generation from the produced audio. Supports batch processing, SRT/VTT/JSON output, language selection, speaker labels, configurable MiniMax voice ID and model. Use when the user asks to turn audio/video into subtitles, generate ASR subtitles, create SRT/VTT, batch transcription, or generate TTS audio with optional subtitles.

Audio To Subtitles

Use this skill to:

  • turn local audio/video files or public media URLs into subtitle files through AI MediaKit;
  • turn text into MiniMax TTS audio;
  • optionally generate subtitles from the TTS audio in the same run.

Workflow

For media inputs:

  1. If the input is a local file, upload it to Cloudflare R2 first.
  2. Submit the public audio_url or video_url to AI MediaKit ASR subtitles.
  3. Poll the async task until it is completed or failed.
  4. Save .json, .srt, and .vtt outputs.

For local video-engine projects:

  1. If the media belongs to a talking-head-hyperframes project under <VIDEO_WORKSPACE>, save subtitles under that project's work/captions/.
  2. Use captions.srt and captions.vtt as the main cleaned files, keep raw ASR backups as captions.raw.srt and captions.raw.vtt, and keep the MediaKit payload as asr-result.json.
  3. Also write captions_aligned.json parsed from the cleaned SRT so video-script, talking-head-hyperframes, and hyperframes-scene-animator share the same locked timeline.
  4. The project-local output convention is:
text
<VIDEO_WORKSPACE>/<project>/work/captions

For text inputs:

  1. Generate audio through MiniMax TTS.
  2. Save the audio file.
  3. If --subtitles is set, upload the audio to R2 and run the same AI MediaKit subtitle workflow.

CLI

bash
# Set this to the directory containing this installed skill (not the current working directory).
SKILL_DIR="<SKILL_ROOT>/audio-to-subtitles"

# Local audio/video: upload to R2, then transcribe
npx -y bun "$SKILL_DIR/scripts/main.ts" audio.mp3 --language zh-CN --out-dir subtitles

# Existing public URL: skip R2 upload
npx -y bun "$SKILL_DIR/scripts/main.ts" "https://example.com/audio.mp3" --out-dir subtitles

# Batch files/URLs
npx -y bun "$SKILL_DIR/scripts/main.ts" a.mp3 b.m4a "https://example.com/video.mp4" --out-dir subtitles

# Batch from manifest
npx -y bun "$SKILL_DIR/scripts/main.ts" --manifest inputs.txt --out-dir subtitles

# Speaker diarization
npx -y bun "$SKILL_DIR/scripts/main.ts" interview.mp3 --speaker --language zh-CN

# Text to audio only
npx -y bun "$SKILL_DIR/scripts/main.ts" --text "你好,欢迎收听。" --voice-id female-shaonv --out-dir audio

# Text file to audio + subtitles
npx -y bun "$SKILL_DIR/scripts/main.ts" --text-file script.md --subtitles --language zh-CN --out-dir audio

Options

OptionDescription
--text <text>Generate MiniMax TTS audio from inline text. Can be repeated.
--text-file <path>Generate MiniMax TTS audio from a text file. Can be repeated.
--subtitlesWith text input, also generate subtitles from the generated audio.
--manifest <path>Batch input manifest. Text files use one input per line; JSON supports an array of strings or objects.
--out-dir <path>Output directory. Default: subtitles.
`--format <allsrt
`--language <cmn-Hans-CNzh-CN
--speakerEnable speaker info and prefix subtitles with speaker labels when returned.
--voice-id <id>MiniMax voice ID. Default: MINIMAX_VOICE_ID, then MINIMAX_TTS_VOICE_ID, then female-shaonv.
--tts-model <model>MiniMax TTS model. Default: MINIMAX_TTS_MODEL, then speech-02-hd.
--tts-speed <n>TTS speed. Default: MINIMAX_TTS_SPEED, then 1.0.
--tts-vol <n>TTS volume. Default: MINIMAX_TTS_VOL, then 1.0.
--tts-pitch <n>TTS pitch. Default: MINIMAX_TTS_PITCH, then 0.
--tts-emotion <emotion>TTS emotion. Default: MINIMAX_TTS_EMOTION, then happy.
`--tts-format <mp3wav
--tts-sample-rate <n>TTS sample rate. Default: MINIMAX_TTS_SAMPLE_RATE, then 32000.
--tts-bitrate <n>TTS bitrate. Default: MINIMAX_TTS_BITRATE, then 128000.
`--media-kind <autoaudio
--r2-prefix <prefix>R2 object key prefix for uploaded local audio/video. Default: R2_AUDIO_KEY_PREFIX, then audio/YYYY-MM-DD.
--concurrency <n>Batch concurrency. Default: 1.
--poll-interval <seconds>Poll interval. Default: 5.
--timeout <seconds>Per-task timeout. Default: 7200.
--jsonPrint machine-readable run summary to stdout.
Show full SKILL.md (234 more words)Show less

Manifest

Text manifests (.txt) are one media input per line.

JSON manifests can mix media and text jobs:

json
[
  "local-audio.mp3",
  { "url": "https://example.com/video.mp4", "mediaKind": "video" },
  { "text": "你好,欢迎收听。", "outputName": "intro", "subtitles": true, "voiceId": "female-shaonv" },
  { "textFile": "script.md", "outputName": "script-audio", "subtitles": true }
]

Environment

The script loads environment files in this order. The nearest .env.r2 is loaded last and overrides global R2 defaults for this vault.

  1. ~/.skills/.env
  2. ~/.baoyu-skills/.env
  3. nearest .env.r2 from current directory upward
  4. nearest .skills/.env from current directory upward
  5. nearest .baoyu-skills/.env from current directory upward

Required for MiniMax TTS text input:

VariableDescription
MINIMAX_API_KEYMiniMax API key.
MINIMAX_TTS_API_KEYAlso accepted.
MINIMAX_API_HOSTOptional. Default: https://api.minimax.io.
MINIMAX_VOICE_IDOptional default voice ID.
MINIMAX_TTS_MODELOptional default model.

Required for MediaKit subtitle generation:

VariableDescription
MEDIAKIT_API_KEYAI MediaKit API key.
AI_MEDIAKIT_API_KEYAlso accepted.
VOLCENGINE_MEDIAKIT_API_KEYAlso accepted.

Required for local-file uploads:

VariableDescription
R2_ACCESS_KEY_IDCloudflare R2 access key.
R2_SECRET_ACCESS_KEYCloudflare R2 secret key.
R2_ACCOUNT_IDCloudflare account ID.
R2_BUCKETR2 bucket name.
R2_PUBLIC_BASE_URLPublic base URL. R2_PUBLIC_URL is also accepted.
R2_AUDIO_KEY_PREFIXOptional R2 prefix for uploaded audio/video. Default: audio/YYYY-MM-DD.

Notes

  • AI MediaKit only accepts public HTTP/HTTPS media URLs. Local files must be uploaded first.
  • TTS audio only needs MiniMax API credentials.
  • TTS + subtitles also needs MediaKit and R2 credentials.
  • Supported API languages are currently Simplified Chinese and English.
  • Single media duration should not exceed the AI MediaKit limit of 3 hours.
  • Do not print API keys or commit real .env.r2 values.
  • Do not leave finished video subtitles outside the video project. Copy or write them into the matching <VIDEO_WORKSPACE>/<project>/work/captions/ directory.

© bozhouDev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/audio-to-subtitles of bozhouDev/video-skills-toolkit.

  • SKILL.md
  • scripts/main.ts
  • scripts/r2.ts

Open the folder on GitHubat commit 4766a16

Compare with similar skills

Audio To Subtitles next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Audio To Subtitles compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Audio To Subtitles this skillbozhouDev/video-skills-toolkit150—~1.9kAutomated safety check: NotesMIT
Hyperframes Mediachmonitor/chmonitor3011 repos~2.8kAutomated safety check: NotesGPL-3.0
Hyperframes CLInateherkai/hyperframes-student-kit1.3k3 repos~1.2kAutomated safety check: PassCustom licence
Content To Videoarchitectds/modeldock117—~2.4kAutomated safety check: PassApache-2.0
Vox ExplainerCK42BB/vox-explainer-skill109—~2.6kAutomated safety check: PassMIT
Hyperframes CLIcoleam00/hyperframes-ai-video-generation149—~1.6kAutomated safety check: PassNone

Similar skills

  • Hyperframes Media

    chmonitor/chmonitor

    Audio and media assets for HyperFrames compositions, produced by one shared audio engine (scripts/audio.mjs) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound…

    301 GitHub starsUsed in 1 repo~2.8k tokens
    Media & CreativeAuto-check: notes
  • Hyperframes CLI

    nateherkai/hyperframes-student-kit

    HyperFrames CLI tool — hyperframes init, lint, preview, render, transcribe, tts, doctor, browser, info, upgrade, compositions, docs, benchmark.

    1.3k GitHub starsUsed in 3 repos~1.2k tokens
    Media & CreativeAuto-check passed
  • Content To Video

    architectds/modeldock

    Turn arbitrary source content (README, article, story, slides, deck, data/report, product description, tutorial text, audio/transcript, or a bare topic) into a finished, high-quality MP4 video.

    117 GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Vox Explainer

    CK42BB/vox-explainer-skill

    End-to-end pipeline for producing Vox-style explainer videos from a single topic prompt.

    109 GitHub stars~2.6k tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed
  • Hyperframes CLI

    coleam00/hyperframes-ai-video-generation

    HyperFrames CLI tool — hyperframes init, lint, inspect, preview, render, transcribe, tts, doctor, browser, info, upgrade, compositions, docs, benchmark.

    149 GitHub stars~1.6k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Videohub Story Editor

    cacity/VideoHub

    把长视频或已有字幕转成有完整叙事的几分钟短片。先基于原文字幕和画面证据理解、选段与重排,再对最终时间轴重新翻译和可选润色;既可输出保留原声的双语字幕版,也可把原声降到 30% 并用 MiniMax 或豆包 TTS 生成影视解说、短剧混剪、播客串讲或知识解读版。已有项目可进入本地五轨时间线继续调整切点、旁白、原声窗口、字幕、音量和转场,并按修订版本渲染。用于“把长视频讲成短故事”“按字幕自动剪辑”…

    168 GitHub stars~2.3k tokensUpdated 9 days ago
    Media & CreativeAuto-check: notes

More from bozhouDev/video-skills-toolkit

All 12 skills in this repo
  • Media To Transcript

    bozhouDev/video-skills-toolkit

    Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0.

    150 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check: notes
  • Talking Head Hyperframes

    bozhouDev/video-skills-toolkit

    为 HyperFrames 口播或旁白项目创建、修复并验证固定舞台,锁定数字人 PIP 的区域、裁切、人物安全区和不透明背景,归档输入,生成 manifest 与 template handoff,并在就绪后按“字幕驱动的全镜头静态审核→动效”门禁路由到 hyperframes-scene-animator。适用于“新建 HyperFrames…

    150 GitHub stars~927 tokensUpdated 2 mo ago
    Auto-check passed
  • Music

    bozhouDev/video-skills-toolkit

    Generate music using ElevenLabs Music API. An agent skill from bozhouDev/video-skills-toolkit.

    150 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Viral Video Benchmark

    bozhouDev/video-skills-toolkit

    判断、扫描、拆解并归档抖音视频、小红书图文或小红书视频。实时读取用户同平台粉丝数并划分主对标池/跨级灵感池,用已登录浏览器读取目标作品和作者主页公开指标,再用确定性代码判定普通、小爆、爆款、现象级并扫描作者近 20 条候选;只对用户选中的爆款和现象级先构建可追溯证据包,再调用子 Agent…

    150 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Douyin Cover

    bozhouDev/video-skills-toolkit

    生成抖音、视频号、小红书等短视频封面图、视频标题图和合集封面,也能诊断和改版已有封面。用户说做封面、生成封面、抖音封面、视频封面、标题图、合集封面、3:4、4:3、1:1、短视频首图、动态封面首帧、给这期视频做图、这封面为什么没人点、帮我改封面、封面点击率怎么提升、诊断封面时都应使用。小白学AI系列封面除外:遇到“小白学AI封面/小白学AI第N集封面”时优先使用…

    150 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Minimax Voice Director

    bozhouDev/video-skills-toolkit

    用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。

    150 GitHub stars~667 tokensUpdated 2 mo ago
    Auto-check: notes

Questions about Audio To Subtitles

What does Audio To Subtitles do?

Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API. Audio To Subtitles is an agent skill from bozhouDev/video-skills-toolkit. Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API.

When should I use Audio To Subtitles?

Audio To Subtitles fits situations like: the user asks to turn audio/video into subtitles; generate ASR subtitles; batch transcription; generate TTS audio with optional subtitles.

How do I install Audio To Subtitles in Claude Code?

Run `npx skills add bozhouDev/video-skills-toolkit --skill audio-to-subtitles -a claude-code`. Or copy the skill folder (skills/audio-to-subtitles in bozhouDev/video-skills-toolkit) into .claude/skills/audio-to-subtitles in your project. Claude Code loads it when a task matches its description.

How do I install Audio To Subtitles in Codex?

Run `npx skills add bozhouDev/video-skills-toolkit --skill audio-to-subtitles -a codex`. Or copy the skill folder (skills/audio-to-subtitles in bozhouDev/video-skills-toolkit) into .agents/skills/audio-to-subtitles in your project. Codex loads it when a task matches its description.

Can I use Audio To Subtitles in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bozhouDev/video-skills-toolkit --skill audio-to-subtitles -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/audio-to-subtitles, .gemini/skills/audio-to-subtitles, .github/skills/audio-to-subtitles and .opencode/skills/audio-to-subtitles in your project.

What does Audio To Subtitles need to run?

Going by SKILL.md and its folder, Audio To Subtitles needs TypeScript for the scripts in its folder, the command-line tools its instructions call (npx) and credentials named MINIMAX_API_KEY, MINIMAX_TTS_API_KEY, MEDIAKIT_API_KEY and AI_MEDIAKIT_API_KEY. Our summary lists: Node.js; A credential in MINIMAX_API_KEY; A credential in MINIMAX_TTS_API_KEY.

Does Audio To Subtitles access the network?

SKILL.md names 1 domain. In commands or code: api.minimax.io; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Audio To Subtitles safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Audio To Subtitles use?

Audio To Subtitles is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Audio To Subtitles use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Audio To Subtitles?

Skills that share tags, products or a category with Audio To Subtitles: Hyperframes Media (chmonitor/chmonitor, 301 stars), Hyperframes CLI (nateherkai/hyperframes-student-kit, 1.3k stars), Content To Video (architectds/modeldock, 117 stars) and Vox Explainer (CK42BB/vox-explainer-skill, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Audio To Subtitles?

bozhouDev (a GitHub user) maintains it in bozhouDev/video-skills-toolkit, which has 150 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on July 27, 2026.

Source: bozhouDev/video-skills-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.