Transcribe Md
hrescak/transcribe-md
Record and transcribe audio to a markdown file using whisper.cpp (mic + system audio)
Audio/video transcription using OpenAI Whisper. An agent skill from MadAppGang/claude-code.
$ npx skills add MadAppGang/claude-code --skill transcription -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install MadAppGang/claude-code transcription --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/MadAppGang/claude-code.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/video-editing/skills/transcription .claude/skills/transcription && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "transcription" agent skill from https://github.com/MadAppGang/claude-code/tree/main/plugins/video-editing/skills/transcription into .claude/skills/transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/MadAppGang/claude-code/tree/main/plugins/video-editing/skills/transcriptionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add MadAppGang/claude-code --skill transcription -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install MadAppGang/claude-code transcription --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MadAppGang/claude-code.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/video-editing/skills/transcription .agents/skills/transcription && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "transcription" agent skill from https://github.com/MadAppGang/claude-code/tree/main/plugins/video-editing/skills/transcription into .agents/skills/transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add MadAppGang/claude-code --skill transcription -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install MadAppGang/claude-code transcription --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MadAppGang/claude-code.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/video-editing/skills/transcription .cursor/skills/transcription && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "transcription" agent skill from https://github.com/MadAppGang/claude-code/tree/main/plugins/video-editing/skills/transcription into .cursor/skills/transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/MadAppGang/claude-code.git --path plugins/video-editing/skills/transcription--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add MadAppGang/claude-code --skill transcription -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install MadAppGang/claude-code transcription --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MadAppGang/claude-code.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/video-editing/skills/transcription .gemini/skills/transcription && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "transcription" agent skill from https://github.com/MadAppGang/claude-code/tree/main/plugins/video-editing/skills/transcription into .gemini/skills/transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install MadAppGang/claude-code transcriptionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add MadAppGang/claude-code --skill transcription -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/MadAppGang/claude-code.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/video-editing/skills/transcription .github/skills/transcription && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "transcription" agent skill from https://github.com/MadAppGang/claude-code/tree/main/plugins/video-editing/skills/transcription into .github/skills/transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add MadAppGang/claude-code --skill transcription -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install MadAppGang/claude-code transcription --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/MadAppGang/claude-code.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/video-editing/skills/transcription .opencode/skills/transcription && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "transcription" agent skill from https://github.com/MadAppGang/claude-code/tree/main/plugins/video-editing/skills/transcription into .opencode/skills/transcription/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "transcription", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
transcriptionAudio/video transcription using OpenAI Whisper. An agent skill from MadAppGang/claude-code.
Transcription is an agent skill from MadAppGang/claude-code. Audio/video transcription using OpenAI Whisper. Covers installation, model selection, transcript formats (SRT, VTT, JSON), timing synchronization, and speaker diarization. Use when transcribing media or generating subtitles.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative, covering Transcription. It works with Whisper. The repository describes itself as: claude code plugins marketplace. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 6097ad4. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
whisperffmpegpipbrewgitmakeffprobeFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Transcription loads about 1.7k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 188 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from MadAppGang/claude-code at commit 6097ad4, republished under its MIT licence (© MadAppGang). 188 words, ~1,674 tokens.
.claude/skills/transcription/SKILL.md (or your agent's skills folder).plugin: video-editing updated: 2026-01-20
Production-ready patterns for audio/video transcription using OpenAI Whisper.
Option 1: OpenAI Whisper (Python)
# macOS/Linux/Windows
pip install openai-whisper
# Verify
whisper --helpOption 2: whisper.cpp (C++ - faster)
# macOS
brew install whisper-cpp
# Linux - build from source
git clone https://github.com/ggerganov/whisper.cpp
cd whisper.cpp && make
# Windows - use pre-built binaries or build with cmakeOption 3: Insanely Fast Whisper (GPU accelerated)
pip install insanely-fast-whisper| Model | Size | VRAM | Accuracy | Speed | Use Case |
|---|---|---|---|---|---|
| tiny | 39M | ~1GB | Low | Fastest | Quick previews |
| base | 74M | ~1GB | Medium | Fast | Draft transcripts |
| small | 244M | ~2GB | Good | Medium | General use |
| medium | 769M | ~5GB | Better | Slow | Quality transcripts |
| large-v3 | 1550M | ~10GB | Best | Slowest | Final production |
Recommendation: Start with small for speed/quality balance. Use large-v3 for final delivery.
# Basic transcription (auto-detect language)
whisper audio.mp3 --model small
# Specify language and output format
whisper audio.mp3 --model medium --language en --output_format srt
# Multiple output formats
whisper audio.mp3 --model small --output_format all
# With timestamps and word-level timing
whisper audio.mp3 --model small --word_timestamps True# Download model first
./models/download-ggml-model.sh base.en
# Transcribe
./main -m models/ggml-base.en.bin -f audio.wav -osrt
# With timestamps
./main -m models/ggml-base.en.bin -f audio.wav -ocsv1
00:00:01,000 --> 00:00:04,500
Hello and welcome to this video.
2
00:00:05,000 --> 00:00:08,200
Today we'll discuss video editing.WEBVTT
00:00:01.000 --> 00:00:04.500
Hello and welcome to this video.
00:00:05.000 --> 00:00:08.200
Today we'll discuss video editing.{
"text": "Hello and welcome to this video.",
"segments": [
{
"id": 0,
"start": 1.0,
"end": 4.5,
"text": " Hello and welcome to this video.",
"words": [
{"word": "Hello", "start": 1.0, "end": 1.3},
{"word": "and", "start": 1.4, "end": 1.5},
{"word": "welcome", "start": 1.6, "end": 2.0},
{"word": "to", "start": 2.1, "end": 2.2},
{"word": "this", "start": 2.3, "end": 2.5},
{"word": "video", "start": 2.6, "end": 3.0}
]
}
]
}Before transcribing video, extract audio in optimal format:
# Extract audio as WAV (16kHz, mono - optimal for Whisper)
ffmpeg -i video.mp4 -ar 16000 -ac 1 -c:a pcm_s16le audio.wav
# Extract as high-quality WAV for archival
ffmpeg -i video.mp4 -vn -c:a pcm_s16le audio.wav
# Extract as compressed MP3 (smaller, still works)
ffmpeg -i video.mp4 -vn -c:a libmp3lame -q:a 2 audio.mp3import json
def whisper_to_fcp_timing(whisper_json_path, fps=24):
"""Convert Whisper JSON output to FCP-compatible timing."""
with open(whisper_json_path) as f:
data = json.load(f)
segments = []
for seg in data.get("segments", []):
segments.append({
"start_time": seg["start"],
"end_time": seg["end"],
"start_frame": int(seg["start"] * fps),
"end_frame": int(seg["end"] * fps),
"text": seg["text"].strip(),
"words": seg.get("words", [])
})
return segments# Get exact frame count and duration
ffprobe -v error -count_frames -select_streams v:0 \
-show_entries stream=nb_read_frames,duration,r_frame_rate \
-of json video.mp4For multi-speaker content, use pyannote.audio:
pip install pyannote.audiofrom pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained("pyannote/speaker-diarization@2.1")
diarization = pipeline("audio.wav")
for turn, _, speaker in diarization.itertracks(yield_label=True):
print(f"{turn.start:.1f}s - {turn.end:.1f}s: {speaker}")#!/bin/bash
# Transcribe all videos in directory
MODEL="small"
OUTPUT_DIR="transcripts"
mkdir -p "$OUTPUT_DIR"
for video in *.mp4 *.mov *.avi; do
[[ -f "$video" ]] || continue
base="${video%.*}"
# Extract audio
ffmpeg -i "$video" -ar 16000 -ac 1 -c:a pcm_s16le "/tmp/${base}.wav" -y
# Transcribe
whisper "/tmp/${base}.wav" --model "$MODEL" \
--output_format all \
--output_dir "$OUTPUT_DIR"
# Cleanup temp audio
rm "/tmp/${base}.wav"
echo "Transcribed: $video"
doneffmpeg -i noisy_audio.wav -af "highpass=f=200,lowpass=f=3000,afftdn=nf=-25" clean_audio.wavwhisper audio.mp3 --language en --model mediumwhisper audio.mp3 --initial_prompt "Technical discussion about video editing software."whisper audio.mp3 --model large-v3 --device cuda# Split audio into 10-minute chunks
# Transcribe each chunk
# Merge results with time offset adjustment# Validate audio file before transcription
validate_audio() {
local file="$1"
if ffprobe -v error -select_streams a:0 -show_entries stream=codec_type -of csv=p=0 "$file" 2>/dev/null | grep -q "audio"; then
return 0
else
echo "Error: No audio stream found in $file"
return 1
fi
}
# Check Whisper installation
check_whisper() {
if command -v whisper &> /dev/null; then
echo "Whisper available"
return 0
else
echo "Error: Whisper not installed. Run: pip install openai-whisper"
return 1
fi
}© MadAppGang, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/video-editing/skills/transcription of MadAppGang/claude-code.
Open the folder on GitHubat commit 6097ad4
Transcription next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Transcription this skillMadAppGang/claude-code | 285 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Transcribe Mdhrescak/transcribe-md | 104 | — | ~474 | Automated safety check: Notes | MIT | |
| Interview Transcriptionjamditis/claude-skills-journalism | 416 | — | ~3.7k | Automated safety check: Pass | MIT | |
| Wjs Transcribing Audiojianshuo/claude-skills | 131 | — | ~4.4k | Automated safety check: Notes | MIT | |
| Whisper Transcriptionbenchflow-ai/skillsbench | 1.8k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Voice Memo SyncLeoYeAI/openclaw-master-skills | 2.2k | — | ~5.1k | Automated safety check: Pass | MIT |
hrescak/transcribe-md
Record and transcribe audio to a markdown file using whisper.cpp (mic + system audio)
jamditis/claude-skills-journalism
Transcription, recording management, and quote extraction. An agent skill from jamditis/claude-skills-journalism.
jianshuo/claude-skills
A skill your agent uses when the user has audio or video and wants a timestamped transcript (SRT) in the source language.
benchflow-ai/skillsbench
Transcribe audio/video to text with word-level timestamps using OpenAI Whisper.
LeoYeAI/openclaw-master-skills
Sync, transcribe, and intelligently organize voice memos, audio/video files, and URLs.
AlexAI-MCP/hermes-CCC
OpenAI Whisper for speech recognition and transcription — local inference, multiple model sizes, language detection, and subtitle generation.
MadAppGang/claude-code
Content brief template and creation methodology for SEO-optimized content.
MadAppGang/claude-code
A skill your agent uses when detecting project technology stack from files/configs/directory structure, auto-loading framework-specific skills, or analyzing multi-stack fullstack projects (e.g…
MadAppGang/claude-code
On-page SEO optimization techniques including keyword density, meta tags, heading structure, and readability.
MadAppGang/claude-code
Techniques for expanding seed keywords and clustering by topic and intent.
MadAppGang/claude-code
SERP analysis techniques for intent classification, feature identification, and competitive intelligence.
MadAppGang/claude-code
A skill your agent uses when deciding whether to launch an agent, selecting which agent to use, or coordinating multiple agents.
Works with
Categories
Audio/video transcription using OpenAI Whisper. An agent skill from MadAppGang/claude-code. Transcription is an agent skill from MadAppGang/claude-code. Audio/video transcription using OpenAI Whisper.
Transcription fits situations like: transcribing media; generating subtitles.
Run `npx skills add MadAppGang/claude-code --skill transcription -a claude-code`. Or copy the skill folder (plugins/video-editing/skills/transcription in MadAppGang/claude-code) into .claude/skills/transcription in your project. Claude Code loads it when a task matches its description.
Run `npx skills add MadAppGang/claude-code --skill transcription -a codex`. Or copy the skill folder (plugins/video-editing/skills/transcription in MadAppGang/claude-code) into .agents/skills/transcription in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add MadAppGang/claude-code --skill transcription -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/transcription, .gemini/skills/transcription, .github/skills/transcription and .opencode/skills/transcription in your project.
Going by SKILL.md and its folder, Transcription needs the command-line tools its instructions call (whisper, ffmpeg, pip, brew, git and make). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Transcription is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Transcription: Transcribe Md (hrescak/transcribe-md, 104 stars), Interview Transcription (jamditis/claude-skills-journalism, 416 stars), Wjs Transcribing Audio (jianshuo/claude-skills, 131 stars) and Whisper Transcription (benchflow-ai/skillsbench, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
MadAppGang (a GitHub organization) maintains it in MadAppGang/claude-code, which has 285 GitHub stars. The repository holds 69 skills in this directory. The repository was last updated on March 15, 2026.
Source: MadAppGang/claude-code on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.