Transcribe video and audio files to SRT subtitles with speaker diarization using the Soniox (default) or AssemblyAI API.

MITAuto-check: notesMedia & Creative

Install Video Transcription

skills CLI
$ npx skills add BlackBeltTechnology/pi-agent-dashboard --skill video-transcription -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BlackBeltTechnology/pi-agent-dashboard video-transcription --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BlackBeltTechnology/pi-agent-dashboard.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/video-transcription/.pi/skills/video-transcription .claude/skills/video-transcription && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-transcription
GitHub stars
315
Token cost
~1.6k tokens
SKILL.md length
718 words
Files
1
Skills in repo
70
Repo updated
First seen
Licence
MIT

At a glance

Transcribe video and audio files to SRT subtitles with speaker diarization using the Soniox (default) or AssemblyAI API.

  • Works in 4 steps: Run pi-transcribe, passing the user's… → The bin handles everything: file… → Present the bin's summary output to the… → …
  • The user wants to transcribe meetings
  • SKILL.md covers Usage, Parallel processing, Execution and Long recordings, plus 1 more section
  • Needs ASSEMBLY_AI_KEY and SONIOX_API_KEY

What it does

Video Transcription is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Transcribe video and audio files to SRT subtitles with speaker diarization using the Soniox (default) or AssemblyAI API. Use when the user wants to transcribe meetings, videos, or audio files. Supports MKV, MP4, MOV, M4A, MP3. Triggers on "/transcribe", "transcribe my meetings", "transcribe videos in ~/Movies", or any request to convert audio/video to text.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Transcription. The repository describes itself as: Real-time web dashboard for pi coding-agent sessions. Multi-session view, live chat mirroring, integrated terminal, diff viewer, pi-flows execution, and mobile-first remote… The licence is MIT.

When your agent uses it

  • The user wants to transcribe meetings
  • Transcribe my meetings
  • Transcribe videos in ~/Movies
  • Any request to convert audio/video to text

Example prompts

  • “/transcribe”
  • “transcribe my meetings”
  • “transcribe videos in ~/Movies”
  • “/video-transcription”

Requirements

  • A credential in ASSEMBLY_AI_KEY
  • A credential in SONIOX_API_KEY

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Run pi-transcribe, passing the user's directory/file arguments (if any)
  2. The bin handles everything: file discovery, audio extraction, API
  3. Present the bin's summary output to the user (total found, already
  4. If there are failures, report which files failed and the error messages

What it can do on your machine

Read from SKILL.md and the folder at commit 7a2d171. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ASSEMBLY_AI_KEY
    • SONIOX_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Transcription loads about 1.6k tokens when it runs. Until then it costs about 95 tokens; SKILL.md has 718 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~95
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:87
    first, then an optional gitignored `.env` (current directory, then the skill

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from BlackBeltTechnology/pi-agent-dashboard at commit 7a2d171, republished under its MIT licence (© BlackBeltTechnology). 718 words, ~1,555 tokens.

Download SKILL.mdSave it as .claude/skills/video-transcription/SKILL.md (or your agent's skills folder).
name
video-transcription
description
Transcribe video and audio files to SRT subtitles with speaker diarization using the Soniox (default) or AssemblyAI API. Use when the user wants to transcribe meetings, videos, or audio files. Supports MKV, MP4, MOV, M4A, MP3. Triggers on "/transcribe", "transcribe my meetings", "transcribe videos in ~/Movies", or any request to convert audio/video to text.

Video Transcription

Transcribe video and audio files in-place to SRT subtitle format with speaker diarization. Backed by the pi-transcribe CLI (a TypeScript port of the original standalone skill — no Python).

Two backends. Soniox is the default and writes <name>.srt. AssemblyAI is opt-in via TRANSCRIBE_BACKEND=assemblyai, uses the EU endpoint, and writes <name>.diarize.srt — a different suffix, so both can transcribe the same file and you can compare diarization side by side.

TRANSCRIBE_BACKEND=both runs them in one pass: audio is extracted once and fed to both APIs, producing <name>.srt and <name>.diarize.srt together. Use this when the user asks to compare diarization or wants both transcripts.

Usage

Run the pi-transcribe bin, optionally passing a directory or file paths. Output .mp3 (extracted audio) and .srt (subtitles) files are placed alongside the source files.

bash
pi-transcribe [directory | file ...]
  • No argument: scans ~/Movies (default)
  • Single directory: scans the specified directory (e.g. a Google Recorder .m4a export folder)
  • One or more file paths: transcribes exactly those files

Examples:

  • /transcribe — transcribe all untranscribed files in ~/Movies
  • /transcribe /path/to/recordings — transcribe files in a specific directory
  • pi-transcribe "~/Movies/May 28 at 4-04 PM.m4a" "~/Movies/Feb 2 at 5-05 PM.m4a" — transcribe specific files
  • TRANSCRIBE_BACKEND=assemblyai pi-transcribe ~/Movies — use AssemblyAI instead of Soniox (writes .diarize.srt, needs ASSEMBLY_AI_KEY)
  • TRANSCRIBE_BACKEND=both pi-transcribe ~/Movies — run BOTH backends in one pass (writes .srt + .diarize.srt; needs both keys)
  • TRANSCRIBE_BACKEND=assemblyai TRANSCRIBE_LANGUAGE=hu pi-transcribe file.m4a — pin Hungarian instead of auto-detecting
  • MAX_CHUNK_HOURS=4 pi-transcribe ~/Movies — change the long-recording chunk size (default 4.5h Soniox / 9h AssemblyAI)
  • TRANSCRIBE_CONCURRENCY=4 pi-transcribe ~/Movies — change how many files transcribe in parallel (default 8)

Parallel processing

Files transcribe through a bounded worker pool: up to TRANSCRIBE_CONCURRENCY files (default 8) are in flight at once, overlapping the Soniox wait that dominates each file's wall-clock time. Files are dispatched oldest-first but may complete in any order. Set TRANSCRIBE_CONCURRENCY=1 for serial, deterministic behavior. The value is clamped to 1–100 (100 = the Soniox pending-job cap).

Execution

  1. Run pi-transcribe, passing the user's directory/file arguments (if any)
  2. The bin handles everything: file discovery, audio extraction, API transcription, idempotency (skips files that already have a sibling .srt)
  3. Present the bin's summary output to the user (total found, already transcribed, newly transcribed, failed)
  4. If there are failures, report which files failed and the error messages

Long recordings

Each provider enforces a HARD per-request limit on audio duration, independent of file size: Soniox 18000 s / 5 h, AssemblyAI 10 h. Recordings longer than the limit are automatically split into MAX_CHUNK_HOURS-sized chunks (default 4.5 h on Soniox, 9 h on AssemblyAI), transcribed separately, and merged into a single SRT with correct absolute timestamps — full coverage, no truncation. Override the chunk size with the MAX_CHUNK_HOURS env var.

Note: the limit is on duration, not megabytes — a long low-bitrate recording can be small in size yet still exceed the cap, so the guard probes duration via ffprobe.

Show full SKILL.md (257 more words)Show less

Prerequisites

  • ffmpeg (with ffprobe) — declared in this package's pi.tools manifest and resolved through the dashboard tool registry (PATH or the static npm packages: ffmpeg-static for ffmpeg, @ffprobe-installer/ffprobe for ffprobe — both optionalDependencies here). When absent, video files are skipped with a warning; audio-only files still process. pi-dashboard-ensure <package-root>/package.json reports both.
  • The selected backend's API key — SONIOX_API_KEY or ASSEMBLY_AI_KEY, both declared as env probes in pi.tools. Same resolver: environment first, then an optional gitignored .env (current directory, then the skill dir). Only the active backend's key is required. No secret ships in the package; the bin fails fast with a clear message if the key is unresolved.
Environment overrides
VariableDefaultMeaning
TRANSCRIBE_BACKENDsonioxsoniox, assemblyai, both/all, or a comma list. Unknown values fall back to soniox.
SONIOX_API_KEY(required for soniox)Soniox API key.
ASSEMBLY_AI_KEY(required for assemblyai)AssemblyAI API key.
MAX_CHUNK_HOURS4.5 / 9Chunk size for long recordings; default follows the backend's duration cap.
MAX_AUDIO_MB200Reserved size guard; 0 disables.
TRANSCRIBE_CONCURRENCY8Files transcribed in parallel. Clamped to 1–100; 1 = serial.
TRANSCRIBE_SRT_SUFFIXper backendOverride the sibling subtitle suffix (.srt / .diarize.srt).
TRANSCRIBE_LANGUAGE(unset)AssemblyAI only: pin a language (e.g. hu) instead of auto-detecting.
TRANSCRIBE_MAX_SPEAKERS(unset)AssemblyAI only: hard cap on speaker labels.
AssemblyAI backend notes

EU endpoint (api.eu.assemblyai.com) for data residency. Runs speech_models: ["universal-3-5-pro", "universal-2"] with speaker_labels and language_detection on: universal-3-5-pro natively covers 18 languages and anything outside that set (Hungarian included) automatically falls back to universal-2 (99 languages). Speaker letters A/B/C are normalised to Speaker 1/Speaker 2/Speaker 3, so the SRT format is identical to Soniox's.

© BlackBeltTechnology, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/video-transcription/.pi/skills/video-transcription of BlackBeltTechnology/pi-agent-dashboard.

Open the folder on GitHubat commit 7a2d171

Compare with similar skills

Video Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Transcription compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Transcription this skillBlackBeltTechnology/pi-agent-dashboard315—~1.6kAutomated safety check: NotesMIT
HyperFrames Media Useheygen-com/hyperframes60k—~2.4kAutomated safety check: PassApache-2.0
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.6k—~2.4kAutomated safety check: PassMIT
Edu Chem Videowy51ai/edulab1.4k—~2.1kAutomated safety check: NotesApache-2.0
Transcription Memory ReconstructionNxcoreAI/EverRoom3k—~714Automated safety check: PassCustom licence
Edu Math Videowy51ai/edulab1.4k—~2.5kAutomated safety check: NotesApache-2.0

Similar skills

  • HyperFrames Media Use

    heygen-com/hyperframes

    Finds, generates and edits media for HyperFrames video projects: music, sound effects, images, icons, logos, voiceovers, captions and color grades.

    60k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.6k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Edu Chem Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a chemistry problem (化学题: 氧化还原配平 双线桥 电子守恒, 物质的量计算, 化学平衡 三段式 平衡常数 转化率 反应速率, 离子反应, 电化学, 溶液 滴定…

    1.4k GitHub stars~2.1k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Reconstruct a complete, searchable memory from an untrusted meeting or conversation transcript.

    3k GitHub stars~714 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Edu Math Video

    wy51ai/edulab

    A skill your agent uses when asked to make an explainer / walkthrough video (讲解视频、解题视频、例题精讲、微课) for a math problem (数学题, geometry, algebra, functions, motion/行程 problems), from a problem screenshot…

    1.4k GitHub stars~2.5k tokensUpdated yesterday
    Media & CreativeAuto-check: notes
  • Transcribe

    JetBrains/skills

    Official

    Transcribe audio files to text with optional diarization and known-speaker hints.

    366 GitHub starsUsed in 4 repos~776 tokens
    Media & CreativeAuto-check passed

More from BlackBeltTechnology/pi-agent-dashboard

All 70 skills in this repo
  • Browser

    BlackBeltTechnology/pi-agent-dashboard

    Browser automation via the agent-browser CLI. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    316 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • CI Troubleshoot

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose failed GitHub Actions runs for pi-agent-dashboard: the 11-file workflow taxonomy, affected-test selection, the release pipeline, known failure modes, and how to read gh run logs and…

    316 GitHub stars~3.5k tokensUpdated today
    Auto-check passed
  • Debug Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Diagnose problems in the running pi-agent-dashboard system: server.log, /api/health, bridge WebSocket connectivity, vitest triage, known-issue FAQ entries.

    316 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Implement

    BlackBeltTechnology/pi-agent-dashboard

    Disciplined implementation in pi-agent-dashboard: the rebuild matrix (extension→reload, server→restart, client→build+restart, openspec-apply→full rebuild) plus the project's code discipline rules.

    316 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Pi Dashboard

    BlackBeltTechnology/pi-agent-dashboard

    Monitor and control the pi-dashboard server. An agent skill from BlackBeltTechnology/pi-agent-dashboard.

    316 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Session To Guideline

    BlackBeltTechnology/pi-agent-dashboard

    Turn a pi session into a Markdown "how-we-did-it" collaboration guideline: reads the session's JSONL transcript and synthesizes a reusable playbook of which prompts worked, what had to be steered…

    316 GitHub stars~3.2k tokensUpdated today
    Auto-check passed

Questions about Video Transcription

What does Video Transcription do?

Transcribe video and audio files to SRT subtitles with speaker diarization using the Soniox (default) or AssemblyAI API. Video Transcription is an agent skill from BlackBeltTechnology/pi-agent-dashboard. Transcribe video and audio files to SRT subtitles with speaker diarization using the Soniox (default) or AssemblyAI API.

When should I use Video Transcription?

Video Transcription fits situations like: the user wants to transcribe meetings; transcribe my meetings; transcribe videos in ~/Movies; any request to convert audio/video to text.

How do I install Video Transcription in Claude Code?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill video-transcription -a claude-code`. Or copy the skill folder (packages/video-transcription/.pi/skills/video-transcription in BlackBeltTechnology/pi-agent-dashboard) into .claude/skills/video-transcription in your project. Claude Code loads it when a task matches its description.

How do I install Video Transcription in Codex?

Run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill video-transcription -a codex`. Or copy the skill folder (packages/video-transcription/.pi/skills/video-transcription in BlackBeltTechnology/pi-agent-dashboard) into .agents/skills/video-transcription in your project. Codex loads it when a task matches its description.

Can I use Video Transcription in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BlackBeltTechnology/pi-agent-dashboard --skill video-transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-transcription, .gemini/skills/video-transcription, .github/skills/video-transcription and .opencode/skills/video-transcription in your project.

What does Video Transcription need to run?

Going by SKILL.md and its folder, Video Transcription needs credentials named ASSEMBLY_AI_KEY and SONIOX_API_KEY. Our summary lists: A credential in ASSEMBLY_AI_KEY; A credential in SONIOX_API_KEY.

Does Video Transcription access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Video Transcription safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Video Transcription use?

Video Transcription is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Transcription use?

About 1.6k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Transcription?

Skills that share tags, products or a category with Video Transcription: HyperFrames Media Use (heygen-com/hyperframes, 60k stars), Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.6k stars), Edu Chem Video (wy51ai/edulab, 1.4k stars) and Transcription Memory Reconstruction (NxcoreAI/EverRoom, 3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Transcription?

BlackBeltTechnology (a GitHub organization) maintains it in BlackBeltTechnology/pi-agent-dashboard, which has 315 GitHub stars. The repository holds 70 skills in this directory. The repository was last updated on October 10, 2026.

Source: BlackBeltTechnology/pi-agent-dashboard on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.