Agent skill

Media To Transcript

by bozhouDev in bozhouDev/video-skills-toolkit

Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0.

MITAuto-check: notesMedia & Creative

Install Media To Transcript

skills CLI
$ npx skills add bozhouDev/video-skills-toolkit --skill media-to-transcript -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install bozhouDev/video-skills-toolkit media-to-transcript --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/bozhouDev/video-skills-toolkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/media-to-transcript .claude/skills/media-to-transcript && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
media-to-transcript
GitHub stars
150
Token cost
~1.8k tokens
SKILL.md length
614 words
Files
6 (incl. scripts, references)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0.

  • Works in 2 steps: Run the doctor if this is the first run,… → Run the pipeline from the notes vault root
  • The user asks for 逐字稿
  • SKILL.md covers Workflow, CLI, Outputs and Options, plus 2 more sections
  • Runs Python scripts from its folder; reaches openspeech.bytedance.com and bilibili.com; needs VOLCENGINE_SPEECH_API_KEY and VOLCENGINE_AUC_API_KEY

What it does

Media To Transcript is an agent skill from bozhouDev/video-skills-toolkit. Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0. Reuses the existing video-transcript downloader for Bilibili, Douyin, Xiaohongshu, YouTube, extracts audio when input is video, uploads local audio to R2, splits media longer than 3 minutes into parallel AUC tasks, submits volc.seedasr.auc with compact source context, then performs AI correction with topic/context/glossary. Use when the user asks for "逐字稿", "提取逐字稿", "视频转文字", "音频转文字", "转写"…

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts and reference files (for example `README.md`, `agents/openai.yaml` and `references/correction-guidelines.md`).

It sits in Media & Creative, covering Transcription, Speech recognition and synthesis and Video and podcast notes. It works with Bilibili, Douyin, Xiaohongshu and YouTube. The repository describes itself as: Video skills toolkit for Remotion talking-head, sketch story, and audio-to-subtitles workflows. The licence is MIT.

When your agent uses it

  • The user asks for 逐字稿
  • Wants a polished transcript

Example prompts

  • “提取视频文案”
  • “/media-to-transcript”

Requirements

  • Python 3
  • A credential in VOLCENGINE_SPEECH_API_KEY
  • A credential in VOLCENGINE_AUC_API_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Run the doctor if this is the first run, after environment changes, or after an error
  2. Run the pipeline from the notes vault root

What it can do on your machine

Read from SKILL.md and the folder at commit 4766a16. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • openspeech.bytedance.com
    • bilibili.com

    Also links to:

    • console.volcengine.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • VOLCENGINE_SPEECH_API_KEY
    • VOLCENGINE_AUC_API_KEY
    • SEEDASR_API_KEY
    • AUC_API_KEY
    • R2_ACCESS_KEY_ID
    • R2_SECRET_ACCESS_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Media To Transcript loads about 1.8k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 145 tokens; SKILL.md has 614 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:103
    key location is your current workspace `.env.r2`, or any environment file loaded by the script:
  • NoteMentions a .env fileSKILL.md:120
    3. Put the key into your workspace `.env.r2` as `VOLCENGINE_SPEECH_API_KEY=...`, or export it in the shell before runnin
  • NoteMentions a .env fileSKILL.md:130
    Do not print API keys or copy `.env.r2` values into chat.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from bozhouDev/video-skills-toolkit at commit 4766a16, republished under its MIT licence (© bozhouDev). 614 words, ~1,776 tokens.

Download SKILL.mdSave it as .claude/skills/media-to-transcript/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
media-to-transcript
description
Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0. Reuses the existing video-transcript downloader for Bilibili, Douyin, Xiaohongshu, YouTube, extracts audio when input is video, uploads local audio to R2, splits media longer than 3 minutes into parallel AUC tasks, submits volc.seedasr.auc with compact source context, then performs AI correction with topic/context/glossary. Use when the user asks for "逐字稿", "提取逐字稿", "视频转文字", "音频转文字", "转写", "听写视频", "提取视频文案", or wants a polished transcript.

Media To Transcript

Use this skill to turn audio or video into a final Markdown transcript. The only ASR backend for this skill is Volcengine recording-file recognition 2.0:

  • submit endpoint: https://openspeech.bytedance.com/api/v3/auc/bigmodel/submit
  • query endpoint: https://openspeech.bytedance.com/api/v3/auc/bigmodel/query
  • resource id: volc.seedasr.auc

Do not use another ASR backend for this skill.

Workflow

  1. Run the doctor if this is the first run, after environment changes, or after an error:
bash
SKILL_DIR="<SKILL_ROOT>/media-to-transcript"
rtk python3 "$SKILL_DIR/scripts/media_to_transcript.py" --doctor
  1. Run the pipeline from the notes vault root:
bash
SKILL_DIR="<SKILL_ROOT>/media-to-transcript"
rtk python3 "$SKILL_DIR/scripts/media_to_transcript.py" "<URL或本地媒体路径>" --topic "视频主题或关键词" --language zh-CN

The script automatically builds AUC context from available media metadata such as title, platform, author, duration, description, and short source URL. It keeps this ASR context under 500 Chinese characters. User-provided --topic, --context, and --context-file are also included in the same 500-character AUC context budget.

For downloaded or local media longer than 180 seconds, the script splits audio into 3-minute chunks, uploads each chunk, submits AUC tasks in parallel, and merges the returned utterances back onto the original timeline.

  1. Read the generated correction-prompt.md and raw-transcript.md.
  2. Apply the AI correction rules in references/correction-guidelines.md.
  3. Save the corrected transcript to the finalTranscriptTarget path printed by the script, and show the full transcript to the user unless they explicitly only asked for a file.

CLI

bash
SKILL_DIR="<SKILL_ROOT>/media-to-transcript"

# Video URL: download video, extract mp3, upload to R2, then run AUC ASR
rtk python3 "$SKILL_DIR/scripts/media_to_transcript.py" "https://www.bilibili.com/video/BVxxx" --topic "AI 编程工具评测"

# Local audio/video
rtk python3 "$SKILL_DIR/scripts/media_to_transcript.py" "/path/to/audio.mp3" --media-kind audio --topic "访谈"

# Add context or terminology hints for ASR and correction
rtk python3 "$SKILL_DIR/scripts/media_to_transcript.py" video.mp4 \
  --topic "Claude Code 和 Codex 使用经验" \
  --context "技术词: Codex, Claude Code, MCP, HyperFrames, R2" \
  --context-file "AI Wiki/raw/some-material.md"

# Enable Volcengine utterance emotion labels when explicitly useful
rtk python3 "$SKILL_DIR/scripts/media_to_transcript.py" video.mp4 --emotion --topic "访谈"

# Convert an existing AUC result JSON into raw transcript + correction prompt
rtk python3 "$SKILL_DIR/scripts/media_to_transcript.py" --from-asr-json asr/auc-result.json --title "视频标题"

Outputs

Each run writes a timestamped directory under outputs/ unless --out-dir is provided:

  • audio/: extracted/converted mp3 when the input is video or unsupported audio
  • audio/chunks/: 3-minute mp3 chunks when media exceeds 180 seconds
  • asr/auc-submit.json or asr/auc-submit-001.json: submitted request metadata, with the API key omitted
  • asr/auc-result.json or asr/auc-result-001.json: raw Volcengine recording-file recognition result
  • asr/auc-result-combined.json: merged AUC result for multi-chunk runs
  • raw-transcript.md: transcript-shaped draft generated from AUC utterances
  • correction-prompt.md: prompt for the AI correction pass
  • transcript.md: target path for the final corrected transcript
  • run-summary.json: machine-readable paths and metadata

The final user-facing artifact is transcript.md. Do not treat auc-result.json as final copy; it is ASR evidence for correction.

Show full SKILL.md (312 more words)Show less

Options

OptionDescription
--title <title>Override detected title.
--topic <text>Topic/theme used during ASR context and correction. Can be repeated.
--context <text>Extra terminology or context for ASR/correction. Can be repeated.
--context-file <path>Read extra context from a local file. Can be repeated.
--language <lang>ASR language. Default: zh-CN.
--speakerRequest speaker clustering when useful for interviews.
--emotionEnable AUC utterance emotion detection. Default: off.
--chunk-seconds <seconds>Split downloaded/local audio longer than this. Default: 180.
--concurrency <n>Parallel AUC chunk tasks. Default: 3.
`--media-kind <autoaudio
`--download <autoalways
--out-dir <path>Output directory. Default: skill outputs/<timestamp>-<title>.
--r2-prefix <prefix>R2 object prefix for uploaded local audio.
--from-asr-json <path>Skip download/ASR and build transcript artifacts from an existing AUC result JSON.
--jsonPrint machine-readable summary to stdout.
--doctorCheck dependencies and required env vars.

API Setup

Preferred key location is your current workspace .env.r2, or any environment file loaded by the script:

bash
VOLCENGINE_SPEECH_API_KEY=你的新版控制台APIKey

Accepted key names:

  • VOLCENGINE_SPEECH_API_KEY (preferred)
  • VOLCENGINE_AUC_API_KEY
  • SEEDASR_API_KEY
  • AUC_API_KEY

If --doctor does not find this API key, use this setup path:

  1. Open 录音文件识别服务开通 and enable “录音文件识别 2.0”.
  2. Open API Key 管理 and create an API Key.
  3. Put the key into your workspace .env.r2 as VOLCENGINE_SPEECH_API_KEY=..., or export it in the shell before running the script.

Local files still need R2 because AUC requires a public audio URL:

  • R2_ACCESS_KEY_ID
  • R2_SECRET_ACCESS_KEY
  • R2_ACCOUNT_ID
  • R2_BUCKET
  • R2_PUBLIC_BASE_URL or R2_PUBLIC_URL

Do not print API keys or copy .env.r2 values into chat.

Correction Rules

Before writing transcript.md, read references/correction-guidelines.md. Core constraints:

  • Correct ASR mistakes using topic/context/glossary and the surrounding transcript.
  • Preserve spoken meaning, order,口语词, repeated words, and uncertainty.
  • Do not summarize, rewrite into article prose, invent missing content, or remove substantive speech.
  • Convert ASR utterance fragments into readable transcript paragraphs with section-level timestamps.
  • Mark unresolved audio as [听不清] instead of guessing.

© bozhouDev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts, references) in skills/media-to-transcript of bozhouDev/video-skills-toolkit.

  • SKILL.md
  • .gitignore
  • README.md
  • agents/openai.yaml
  • references/correction-guidelines.md
  • scripts/media_to_transcript.py

Open the folder on GitHubat commit 4766a16

Compare with similar skills

Media To Transcript next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Media To Transcript compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Media To Transcript this skillbozhouDev/video-skills-toolkit150—~1.8kAutomated safety check: NotesMIT
Video SummaryLeoYeAI/openclaw-master-skills2.2k—~4.2kAutomated safety check: PassMIT
Video To NotesKIRVO-REPORTING/video-to-notes105—~1.5kAutomated safety check: PassMIT
Ra Video DownloadPluviobyte/rnskill1.6k—~861Automated safety check: NotesCustom licence
Video Downloaderkangarooking/kangarooking-skills657—~8.3kAutomated safety check: PassNone
Video Podcast MakerAgents365-ai/video-podcast-maker1.7k—~4.9kAutomated safety check: PassMIT

Similar skills

  • Video Summary

    LeoYeAI/openclaw-master-skills

    Video summarization for Bilibili, Xiaohongshu, Douyin, and YouTube.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Video To Notes

    KIRVO-REPORTING/video-to-notes

    Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.

    105 GitHub stars~1.5k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Ra Video Download

    Pluviobyte/rnskill

    Download source video or audio from Douyin, YouTube, Bilibili, Twitter/X, Xiaohongshu, and other yt-dlp-supported URLs into the content-creation workspace.

    1.6k GitHub stars~861 tokensUpdated 17 days ago
    Media & CreativeAuto-check: notes
  • Video Downloader

    kangarooking/kangarooking-skills

    Download or open videos and recover platform captions, audio transcripts, keyframes, screen text, visual facts, and editing observations as a plain multimodaltranscript.md.

    657 GitHub stars~8.3k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video Podcast Maker

    Agents365-ai/video-podcast-maker

    A skill your agent uses when the user gives a topic and wants an automated topic-driven narrated explainer, podcast, or knowledge-summary video (Bilibili / YouTube / Xiaohongshu / Douyin / WeChat…

    1.7k GitHub stars~4.9k tokensUpdated 7 days ago
    Media & CreativeAuto-check passed
  • Video To Subtitle Summary

    imlewc/video-to-subtitle-summary-skill

    A skill your agent uses when user provides a short video platform URL or local video/audio file and wants subtitles/AI summary, or when user asks to list their own AI Douyin historical tasks.

    212 GitHub stars~4.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes

More from bozhouDev/video-skills-toolkit

All 12 skills in this repo
  • Audio To Subtitles

    bozhouDev/video-skills-toolkit

    Convert local audio/video files or public media URLs into subtitle files by uploading local files to Cloudflare R2 and calling Volcengine AI MediaKit ASR subtitles API.

    150 GitHub stars~1.9k tokensUpdated 2 mo ago
    Auto-check: notes
  • Talking Head Hyperframes

    bozhouDev/video-skills-toolkit

    为 HyperFrames 口播或旁白项目创建、修复并验证固定舞台,锁定数字人 PIP 的区域、裁切、人物安全区和不透明背景,归档输入,生成 manifest 与 template handoff,并在就绪后按“字幕驱动的全镜头静态审核→动效”门禁路由到 hyperframes-scene-animator。适用于“新建 HyperFrames…

    150 GitHub stars~927 tokensUpdated 2 mo ago
    Auto-check passed
  • Music

    bozhouDev/video-skills-toolkit

    Generate music using ElevenLabs Music API. An agent skill from bozhouDev/video-skills-toolkit.

    150 GitHub stars~3.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Viral Video Benchmark

    bozhouDev/video-skills-toolkit

    判断、扫描、拆解并归档抖音视频、小红书图文或小红书视频。实时读取用户同平台粉丝数并划分主对标池/跨级灵感池,用已登录浏览器读取目标作品和作者主页公开指标,再用确定性代码判定普通、小爆、爆款、现象级并扫描作者近 20 条候选;只对用户选中的爆款和现象级先构建可追溯证据包,再调用子 Agent…

    150 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Douyin Cover

    bozhouDev/video-skills-toolkit

    生成抖音、视频号、小红书等短视频封面图、视频标题图和合集封面,也能诊断和改版已有封面。用户说做封面、生成封面、抖音封面、视频封面、标题图、合集封面、3:4、4:3、1:1、短视频首图、动态封面首帧、给这期视频做图、这封面为什么没人点、帮我改封面、封面点击率怎么提升、诊断封面时都应使用。小白学AI系列封面除外:遇到“小白学AI封面/小白学AI第N集封面”时优先使用…

    150 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • Minimax Voice Director

    bozhouDev/video-skills-toolkit

    用 MiniMax 云端为视频制作可审批的声音导演稿,再生成、挑选和验收人声,最后以定稿音频产生字幕。用于用户明确选择 MiniMax 配音、继续已有 MiniMax 视频配音项目,或明确请求 MiniMax Voice ID/克隆/设计。泛指本地 TTS 或 IndexTTS 不使用本 skill;音乐、BGM、歌曲使用同级 music Skill。

    150 GitHub stars~667 tokensUpdated 2 mo ago
    Auto-check: notes

Questions about Media To Transcript

What does Media To Transcript do?

Convert audio/video URLs or local media into corrected Markdown transcripts through Volcengine recording-file ASR 2.0. Media To Transcript is an agent skill from bozhouDev/video-skills-toolkit.0.

When should I use Media To Transcript?

Media To Transcript fits situations like: the user asks for 逐字稿; wants a polished transcript.

How do I install Media To Transcript in Claude Code?

Run `npx skills add bozhouDev/video-skills-toolkit --skill media-to-transcript -a claude-code`. Or copy the skill folder (skills/media-to-transcript in bozhouDev/video-skills-toolkit) into .claude/skills/media-to-transcript in your project. Claude Code loads it when a task matches its description.

How do I install Media To Transcript in Codex?

Run `npx skills add bozhouDev/video-skills-toolkit --skill media-to-transcript -a codex`. Or copy the skill folder (skills/media-to-transcript in bozhouDev/video-skills-toolkit) into .agents/skills/media-to-transcript in your project. Codex loads it when a task matches its description.

Can I use Media To Transcript in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add bozhouDev/video-skills-toolkit --skill media-to-transcript -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/media-to-transcript, .gemini/skills/media-to-transcript, .github/skills/media-to-transcript and .opencode/skills/media-to-transcript in your project.

What does Media To Transcript need to run?

Going by SKILL.md and its folder, Media To Transcript needs Python for the scripts in its folder and credentials named VOLCENGINE_SPEECH_API_KEY, VOLCENGINE_AUC_API_KEY, SEEDASR_API_KEY and AUC_API_KEY. Our summary lists: Python 3; A credential in VOLCENGINE_SPEECH_API_KEY; A credential in VOLCENGINE_AUC_API_KEY.

Does Media To Transcript access the network?

SKILL.md names 3 domains. In commands or code: openspeech.bytedance.com and bilibili.com; the agent is likely to contact these when it follows the instructions. As links in the text: console.volcengine.com. This is read from the text; nothing was executed.

Is Media To Transcript safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Media To Transcript use?

Media To Transcript is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Media To Transcript use?

About 1.8k tokens (SKILL.md is roughly 7.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 595 tokens, read only when the agent opens those files.

What are the alternatives to Media To Transcript?

Skills that share tags, products or a category with Media To Transcript: Video Summary (LeoYeAI/openclaw-master-skills, 2.2k stars), Video To Notes (KIRVO-REPORTING/video-to-notes, 105 stars), Ra Video Download (Pluviobyte/rnskill, 1.6k stars) and Video Downloader (kangarooking/kangarooking-skills, 657 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Media To Transcript?

bozhouDev (a GitHub user) maintains it in bozhouDev/video-skills-toolkit, which has 150 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on July 27, 2026.

Source: bozhouDev/video-skills-toolkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.