Agent skill

Conference Transcribe

by swyxio in swyxio/skills

Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts.

MITAuto-check passedMedia & Creative

Install Conference Transcribe

skills CLI
$ npx skills add swyxio/skills --skill conference-transcribe -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills conference-transcribe --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/conference-transcribe .claude/skills/conference-transcribe && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
conference-transcribe
GitHub stars
176
Token cost
~2.4k tokens
SKILL.md length
676 words
Files
1
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts.

  • Works in 8 steps: Check Prerequisites → Get Video Metadata and Captions → Parse Chapter Timestamps into Talk… → …
  • User says transcribe this conference
  • SKILL.md covers Lessons Learned (from AIE…, Step-by-Step Workflow and Quick Start (Copy-Paste)
  • Calls yt-dlp, ffmpeg and python3; reaches youtube.com and api.groq.com; needs HF_TOKEN and GROQ_API_KEY

What it does

Conference Transcribe is an agent skill from swyxio/skills. Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts. Parses timestamps from the video description to split talks, downloads audio/video, transcribes each segment, then uses an LLM to clean up and format the transcripts with key takeaways and frequent timestamps. Use when user says "transcribe this conference", "split this livestream into talks", "transcribe each talk separately", or provides a YouTube URL of a multi-hour event stream with chapter timestamps.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic…

It sits in Media & Creative, covering Transcription. It works with YouTube. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • User says transcribe this conference
  • Split this livestream into talks
  • Transcribe each talk separately
  • Provides a YouTube URL of a multi-hour event stream with chapter timestamps

Example prompts

  • “transcribe this conference”
  • “split this livestream into talks”
  • “transcribe each talk separately”
  • “/conference-transcribe”

Requirements

  • Python 3
  • A credential in GROQ_API_KEY
  • A credential in ANTHROPIC_API_KEY
  • Compatibility (from SKILL.md): Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic recommended) for cleanup pass.

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Check Prerequisites
  2. Get Video Metadata and Captions
  3. Parse Chapter Timestamps into Talk Manifest
  4. Build Raw Transcripts (Caption Path)
  5. (alt): Build Raw Transcripts (Whisper/API Path)
  6. Download Video Clips (Optional, Parallel)
  7. LLM Cleanup Pass
  8. Write Final Output

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • yt-dlp
    • ffmpeg
    • python3
    • uv
    • curl
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com
    • api.groq.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • HF_TOKEN
    • GROQ_API_KEY
    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic recommended) for cleanup pass.

    From compatibility in the SKILL.md frontmatter.

Context cost

Conference Transcribe loads about 2.4k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 676 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 676 words, ~2,423 tokens.

Download SKILL.mdSave it as .claude/skills/conference-transcribe/SKILL.md (or your agent's skills folder).
name
conference-transcribe
description
Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts. Parses timestamps from the video description to split talks, downloads audio/video, transcribes each segment, then uses an LLM to clean up and format the transcripts with key takeaways and frequent timestamps. Use when user says "transcribe this conference", "split this livestream into talks", "transcribe each talk separately", or provides a YouTube URL of a multi-hour event stream with chapter timestamps.
compatibility
Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic recommended) for cleanup pass.
license
MIT
metadata.author
swyxio
metadata.version
1.0
metadata.last-updated
2026-04-10
metadata.primary-tools
yt-dlp, ffmpeg, Groq Whisper API, Claude API

Conference Transcribe

Input: a YouTube URL with chapter timestamps or another multi-talk recording.

Transcribe a multi-talk conference livestream into individual, cleaned-up per-talk markdown files with key takeaways and frequent timestamps.

Lessons Learned (from AIE Europe Day 1, April 2026)

This skill was born from a real transcription session. Here's what worked and what didn't:

What Worked
  1. YouTube auto-captions as primary source -- YouTube's English auto-captions (VTT) are surprisingly good and instant. No download wait, no model download, no GPU needed. Prefer this over Whisper for YouTube videos that have captions available.
  2. yt-dlp --write-sub --write-auto-sub --sub-lang en to grab captions directly, then parse the VTT file.
  3. yt-dlp --download-sections to download individual talk clips as MP4 (for archiving/sharing), parallelized with ThreadPoolExecutor(max_workers=2).
  4. Parsing the video description for chapter timestamps -- the description is the ground truth for talk boundaries.
  5. Two-pass transcript pipeline: raw VTT parse -> LLM cleanup. The raw parse handles deduplication and timestamp bucketing; the LLM handles readability, proper nouns, key takeaways, and formatting.
  6. Opus compression for cloud API uploads: ffmpeg -c:a libopus -b:a 32k gets 1 hour of audio down to ~14MB (under Groq/OpenAI's 25MB limit).
  7. metadata.json from yt-dlp --write-info-json gives structured chapter data that's easier to parse than the description text.
What Didn't Work
  1. Local Whisper (any variant) on a fresh machine is painful:
    • pip install fights with macOS externally-managed-environment (PEP 668). Need uv venv or --break-system-packages.
    • mlx-whisper model downloads are huge (~1.5GB for large-v3-turbo) and get throttled without HF_TOKEN.
    • faster-whisper on CPU is very slow for 7+ hours of audio.
    • All told, setting up local Whisper from scratch took longer than just using YouTube captions.
  2. Groq API without a key -- need to ask user upfront.
  3. Parallel Whisper on Apple Silicon -- MLX uses the GPU, so parallel workers contend. Sequential is better for local models.
  4. Large WAV files for API upload -- 16kHz mono WAV is ~1.9MB/min. A 30-min talk is 57MB, over the 25MB API limit. Must compress to opus/ogg first.
Decision Tree: Which Transcription Backend?
Has YouTube captions? (check with yt-dlp --list-subs)
  YES -> Use YouTube captions (fastest, free, no setup)
  NO  -> Has Groq/OpenAI API key?
    YES -> Use Groq API (whisper-large-v3-turbo, nearly free, very fast)
    NO  -> Has local Whisper installed?
      YES -> Use faster-whisper or mlx-whisper
      NO  -> Install via uv: `uv venv .venv && source .venv/bin/activate && uv pip install faster-whisper`

Step-by-Step Workflow

Step 0: Check Prerequisites
bash
which yt-dlp || echo "MISSING: brew install yt-dlp"
which ffmpeg || echo "MISSING: brew install ffmpeg"
which jq || echo "MISSING: brew install jq"

# Check for API keys (optional but recommended)
[ -n "$GROQ_API_KEY" ] && echo "Groq: ready" || echo "Groq: not set (needed for Whisper API)"
[ -n "$ANTHROPIC_API_KEY" ] && echo "Anthropic: ready" || echo "Anthropic: not set (needed for cleanup)"
Step 1: Get Video Metadata and Captions
bash
VIDEO_URL="$1"  # YouTube URL from user

# Download metadata + captions (no video)
yt-dlp --write-info-json --skip-download \
  --write-sub --write-auto-sub --sub-lang en \
  -o "media/%(id)s" "$VIDEO_URL"

This produces:

  • media/<id>.info.json -- full metadata including chapters
  • media/<id>.en-orig.vtt or media/<id>.en.vtt -- auto-captions
Step 2: Parse Chapter Timestamps into Talk Manifest

Read the .info.json file and extract chapters. Filter out breaks, untitled segments, etc.

Create a talks.json manifest:

json
[
  {
    "index": 1,
    "title": "Speaker Name: Talk Title",
    "speaker": "Speaker Name",
    "slug": "01-speaker-name",
    "source_chapter_start": "00:24:25",
    "source_chapter_end": "00:42:39",
    "start_seconds": 1465,
    "end_seconds": 2559,
    "duration_seconds": 1094
  }
]

If no chapters in metadata, fall back to parsing the video description for timestamp lines matching patterns like:

  • HH:MM:SS - Speaker Name: Talk Title
  • HH:MM:SS Speaker Name (Company): Description
Show full SKILL.md (308 more words)Show less
Step 3: Build Raw Transcripts (Caption Path)

If YouTube captions are available, parse the VTT file:

  1. Parse all VTT cues (timestamp + text).
  2. For each talk in the manifest, select cues within the talk's time range.
  3. Deduplicate overlapping text -- YouTube VTT cues often repeat words from the previous cue. Use a sliding window word-match approach (see append_without_overlap pattern).
  4. Bucket cues into ~30-second paragraphs for readability.
  5. Write each talk as a raw transcript file with dual timestamps: [HH:MM:SS | +MM:SS] (absolute stream time | relative to talk start).

Output format:

markdown
# Speaker Name: Talk Title

- Source: https://youtube.com/watch?v=ID&t=1465s
- Source range: 00:24:25 - 00:42:39
- Duration: 00:18:14
- Transcript source: YouTube auto-captions

## Timestamped Transcript

[00:24:26 | +00:00:01] Good morning everyone...

[00:24:54 | +00:00:29] Next paragraph of text...
Step 3 (alt): Build Raw Transcripts (Whisper/API Path)

If no YouTube captions, transcribe from audio:

  1. Download audio: yt-dlp -f bestaudio -x --audio-format wav -o "full_audio.%(ext)s" "$VIDEO_URL"
  2. Split into per-talk segments using ffmpeg:
    bash
    ffmpeg -i full_audio.wav -ss "$START" -to "$END" -ac 1 -ar 16000 "segments/${SLUG}.wav"
  3. Compress for API upload:
    bash
    ffmpeg -i "segments/${SLUG}.wav" -ac 1 -ar 16000 -c:a libopus -b:a 32k "segments/${SLUG}.ogg"
  4. If ogg > 25MB, further split into 10-min chunks.
  5. Transcribe via Groq API (preferred) or other backend:
    bash
    curl -s https://api.groq.com/openai/v1/audio/transcriptions \
      -H "Authorization: Bearer $GROQ_API_KEY" \
      -F file="@segments/${SLUG}.ogg" \
      -F model="whisper-large-v3-turbo" \
      -F language="en" \
      -F response_format="verbose_json" \
      -F 'timestamp_granularities[]=segment'
  6. Reassemble chunks with offset correction for absolute timestamps.
  7. Parallelize: 3 concurrent API calls is safe for Groq rate limits.
Step 4: Download Video Clips (Optional, Parallel)

For archiving individual talk videos:

bash
yt-dlp -f 91 \
  --downloader ffmpeg \
  --downloader-args "ffmpeg_i:-allowed_extensions ALL" \
  --download-sections "*${START}-${END}" \
  -o "clips/${SLUG}.%(ext)s" \
  "$VIDEO_URL"

Run with ThreadPoolExecutor(max_workers=2) -- more than 2 concurrent yt-dlp downloads tends to get throttled.

Step 5: LLM Cleanup Pass

Send each raw transcript to Claude (or another LLM) for cleanup. Use 3 concurrent API calls.

System prompt for cleanup:

You are an expert transcript editor. Take this raw auto-caption transcript and produce a clean, readable document.

Rules:
1. KEEP all timestamps in [HH:MM:SS | +MM:SS] format. Include them every 30-60 seconds.
2. Fix transcription errors: proper nouns, technical terms, company names, jargon.
3. Add paragraph breaks at natural topic transitions.
4. Remove filler words (um, uh, like, you know) unless they add meaning.
5. Preserve the speaker's voice -- clean up, don't rewrite.
6. For multi-speaker segments, use [Speaker Name]: format.

Output format:
# Talk Title
**Speaker** -- Role/Company
**Event**: {event name and date}

## Key Points
- 4-8 bullet points of key takeaways

## Timestamped Reading Transcript
[Content with timestamps, paragraphs, and light line-wrapping for readability]
Step 6: Write Final Output

Organize into:

transcripts/
  raw/       -- unedited VTT-parsed or Whisper output
  cleaned/   -- LLM-cleaned markdown
clips/       -- individual talk MP4s (optional)
talks.json   -- manifest with metadata
reports/
  talk-manifest.md  -- summary table

Quick Start (Copy-Paste)

For the common case of a YouTube conference stream with chapters and auto-captions:

bash
# 1. Grab metadata + captions
yt-dlp --write-info-json --skip-download --write-auto-sub --sub-lang en -o "media/%(id)s" "$URL"

# 2. Build talks.json from chapters in info.json
python3 scripts/build_transcripts.py

# 3. Download individual clips (parallel, optional)
python3 scripts/download_clips.py

# 4. Clean up transcripts with LLM
python3 scripts/cleanup_transcripts.py

The build_transcripts.py and cleanup_transcripts.py scripts handle VTT parsing, deduplication, timestamp formatting, and LLM cleanup. See the AIE Europe project for reference implementations.

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in conference-transcribe of swyxio/skills.

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Conference Transcribe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Conference Transcribe compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Conference Transcribe this skillswyxio/skills176—~2.4kAutomated safety check: PassMIT
Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image2.6k—~2.4kAutomated safety check: PassMIT
Video Dataoxylabs/agent-skills875—~1.4kAutomated safety check: PassMIT
Summarizetrpc-group/trpc-agent-go1.9k22 repos~552Automated safety check: PassApache-2.0
Youtube PublishAndonywang123/Epost197—~3.4kAutomated safety check: WarnNone
Youtube Transcribe Skillfeiskyer/codex-settings244—~745Automated safety check: PassMIT

Similar skills

  • Native Subtitle Quote Image

    chengyi-ai/native-subtitle-quote-image

    将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…

    2.6k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Video Data

    oxylabs/agent-skills

    YouTube data extraction API and high-bandwidth proxy downloads.

    875 GitHub stars~1.4k tokensUpdated 11 days ago
    Media & CreativeAuto-check passed
  • Summarize

    trpc-group/trpc-agent-go

    Summarize or extract text/transcripts from URLs, podcasts, and local files (great fallback for “transcribe this YouTube/video”).

    1.9k GitHub starsUsed in 22 repos~552 tokens
    Media & CreativeAuto-check passed
  • Youtube Publish

    Andonywang123/Epost

    Prepare an English YouTube release with local Chinese-to-English translation, subtitles and cover localization, then use a deterministic script connected to dedicated Chrome and YouTube Studio to…

    197 GitHub stars~3.4k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: warnings
  • Youtube Transcribe Skill

    feiskyer/codex-settings

    Extract subtitles or a transcript from a YouTube URL and save normalized timestamped text locally.

    244 GitHub stars~745 tokensUpdated 14 days ago
    Media & CreativeAuto-check passed
  • Watching Videos

    oxbshw/watch-skill

    The user shared a video URL, a YouTube/TikTok/stream link, a local video file, a screen recording, a meeting recording, or a playlist/folder of videos — "watch this", "summarize this video", "what's…

    470 GitHub stars~599 tokensUpdated 26 days ago
    Media & CreativeAuto-check: notes

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    176 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    176 GitHub stars~4.3k tokensUpdated today
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    176 GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    176 GitHub stars~1.5k tokensUpdated today
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    176 GitHub stars~1.8k tokensUpdated today
    Auto-check passed

Works with

Questions about Conference Transcribe

What does Conference Transcribe do?

Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts. Conference Transcribe is an agent skill from swyxio/skills. Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts.

When should I use Conference Transcribe?

Conference Transcribe fits situations like: user says transcribe this conference; split this livestream into talks; transcribe each talk separately; provides a YouTube URL of a multi-hour event stream with chapter timestamps.

How do I install Conference Transcribe in Claude Code?

Run `npx skills add swyxio/skills --skill conference-transcribe -a claude-code`. Or copy the skill folder (conference-transcribe in swyxio/skills) into .claude/skills/conference-transcribe in your project. Claude Code loads it when a task matches its description.

How do I install Conference Transcribe in Codex?

Run `npx skills add swyxio/skills --skill conference-transcribe -a codex`. Or copy the skill folder (conference-transcribe in swyxio/skills) into .agents/skills/conference-transcribe in your project. Codex loads it when a task matches its description.

Can I use Conference Transcribe in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill conference-transcribe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/conference-transcribe, .gemini/skills/conference-transcribe, .github/skills/conference-transcribe and .opencode/skills/conference-transcribe in your project.

What does Conference Transcribe need to run?

Going by SKILL.md and its folder, Conference Transcribe needs the command-line tools its instructions call (yt-dlp, ffmpeg, python3, uv, curl and pip) and credentials named HF_TOKEN, GROQ_API_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in GROQ_API_KEY; A credential in ANTHROPIC_API_KEY. Compatibility (from SKILL.md): Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic recommended) for cleanup pass. .

Does Conference Transcribe access the network?

SKILL.md names 2 domains. In commands or code: youtube.com and api.groq.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Conference Transcribe safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Conference Transcribe use?

Conference Transcribe is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Conference Transcribe use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Conference Transcribe?

Skills that share tags, products or a category with Conference Transcribe: Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.6k stars), Video Data (oxylabs/agent-skills, 875 stars), Summarize (trpc-group/trpc-agent-go, 1.9k stars) and Youtube Publish (Andonywang123/Epost, 197 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Conference Transcribe?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 176 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.