Native Subtitle Quote Image
chengyi-ai/native-subtitle-quote-image
将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…
Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts.
$ npx skills add swyxio/skills --skill conference-transcribe -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install swyxio/skills conference-transcribe --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/conference-transcribe .claude/skills/conference-transcribe && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "conference-transcribe" agent skill from https://github.com/swyxio/skills/tree/main/conference-transcribe into .claude/skills/conference-transcribe/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conference-transcribe", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/swyxio/skills/tree/main/conference-transcribeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add swyxio/skills --skill conference-transcribe -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install swyxio/skills conference-transcribe --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/conference-transcribe .agents/skills/conference-transcribe && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "conference-transcribe" agent skill from https://github.com/swyxio/skills/tree/main/conference-transcribe into .agents/skills/conference-transcribe/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conference-transcribe", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill conference-transcribe -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install swyxio/skills conference-transcribe --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/conference-transcribe .cursor/skills/conference-transcribe && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "conference-transcribe" agent skill from https://github.com/swyxio/skills/tree/main/conference-transcribe into .cursor/skills/conference-transcribe/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conference-transcribe", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/swyxio/skills.git --path conference-transcribe--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add swyxio/skills --skill conference-transcribe -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install swyxio/skills conference-transcribe --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/conference-transcribe .gemini/skills/conference-transcribe && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "conference-transcribe" agent skill from https://github.com/swyxio/skills/tree/main/conference-transcribe into .gemini/skills/conference-transcribe/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conference-transcribe", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install swyxio/skills conference-transcribeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add swyxio/skills --skill conference-transcribe -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/conference-transcribe .github/skills/conference-transcribe && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "conference-transcribe" agent skill from https://github.com/swyxio/skills/tree/main/conference-transcribe into .github/skills/conference-transcribe/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conference-transcribe", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill conference-transcribe -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install swyxio/skills conference-transcribe --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/conference-transcribe .opencode/skills/conference-transcribe && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "conference-transcribe" agent skill from https://github.com/swyxio/skills/tree/main/conference-transcribe into .opencode/skills/conference-transcribe/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "conference-transcribe", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
conference-transcribeTranscribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts.
Conference Transcribe is an agent skill from swyxio/skills. Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts. Parses timestamps from the video description to split talks, downloads audio/video, transcribes each segment, then uses an LLM to clean up and format the transcripts with key takeaways and frequent timestamps. Use when user says "transcribe this conference", "split this livestream into talks", "transcribe each talk separately", or provides a YouTube URL of a multi-hour event stream with chapter timestamps.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts. Compatibility notes: Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic…
It sits in Media & Creative, covering Transcription. It works with YouTube. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
yt-dlpffmpegpython3uvcurlpipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
youtube.comapi.groq.comFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
HF_TOKENGROQ_API_KEYANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic recommended) for cleanup pass.
From compatibility in the SKILL.md frontmatter.
Conference Transcribe loads about 2.4k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 676 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 676 words, ~2,423 tokens.
.claude/skills/conference-transcribe/SKILL.md (or your agent's skills folder).Input: a YouTube URL with chapter timestamps or another multi-talk recording.
Transcribe a multi-talk conference livestream into individual, cleaned-up per-talk markdown files with key takeaways and frequent timestamps.
This skill was born from a real transcription session. Here's what worked and what didn't:
yt-dlp --write-sub --write-auto-sub --sub-lang en to grab captions directly, then parse the VTT file.yt-dlp --download-sections to download individual talk clips as MP4 (for archiving/sharing), parallelized with ThreadPoolExecutor(max_workers=2).ffmpeg -c:a libopus -b:a 32k gets 1 hour of audio down to ~14MB (under Groq/OpenAI's 25MB limit).metadata.json from yt-dlp --write-info-json gives structured chapter data that's easier to parse than the description text.pip install fights with macOS externally-managed-environment (PEP 668). Need uv venv or --break-system-packages.mlx-whisper model downloads are huge (~1.5GB for large-v3-turbo) and get throttled without HF_TOKEN.faster-whisper on CPU is very slow for 7+ hours of audio.Has YouTube captions? (check with yt-dlp --list-subs)
YES -> Use YouTube captions (fastest, free, no setup)
NO -> Has Groq/OpenAI API key?
YES -> Use Groq API (whisper-large-v3-turbo, nearly free, very fast)
NO -> Has local Whisper installed?
YES -> Use faster-whisper or mlx-whisper
NO -> Install via uv: `uv venv .venv && source .venv/bin/activate && uv pip install faster-whisper`which yt-dlp || echo "MISSING: brew install yt-dlp"
which ffmpeg || echo "MISSING: brew install ffmpeg"
which jq || echo "MISSING: brew install jq"
# Check for API keys (optional but recommended)
[ -n "$GROQ_API_KEY" ] && echo "Groq: ready" || echo "Groq: not set (needed for Whisper API)"
[ -n "$ANTHROPIC_API_KEY" ] && echo "Anthropic: ready" || echo "Anthropic: not set (needed for cleanup)"VIDEO_URL="$1" # YouTube URL from user
# Download metadata + captions (no video)
yt-dlp --write-info-json --skip-download \
--write-sub --write-auto-sub --sub-lang en \
-o "media/%(id)s" "$VIDEO_URL"This produces:
media/<id>.info.json -- full metadata including chaptersmedia/<id>.en-orig.vtt or media/<id>.en.vtt -- auto-captionsRead the .info.json file and extract chapters. Filter out breaks, untitled segments, etc.
Create a talks.json manifest:
[
{
"index": 1,
"title": "Speaker Name: Talk Title",
"speaker": "Speaker Name",
"slug": "01-speaker-name",
"source_chapter_start": "00:24:25",
"source_chapter_end": "00:42:39",
"start_seconds": 1465,
"end_seconds": 2559,
"duration_seconds": 1094
}
]If no chapters in metadata, fall back to parsing the video description for timestamp lines matching patterns like:
HH:MM:SS - Speaker Name: Talk TitleHH:MM:SS Speaker Name (Company): DescriptionIf YouTube captions are available, parse the VTT file:
append_without_overlap pattern).[HH:MM:SS | +MM:SS] (absolute stream time | relative to talk start).Output format:
# Speaker Name: Talk Title
- Source: https://youtube.com/watch?v=ID&t=1465s
- Source range: 00:24:25 - 00:42:39
- Duration: 00:18:14
- Transcript source: YouTube auto-captions
## Timestamped Transcript
[00:24:26 | +00:00:01] Good morning everyone...
[00:24:54 | +00:00:29] Next paragraph of text...If no YouTube captions, transcribe from audio:
yt-dlp -f bestaudio -x --audio-format wav -o "full_audio.%(ext)s" "$VIDEO_URL"ffmpeg -i full_audio.wav -ss "$START" -to "$END" -ac 1 -ar 16000 "segments/${SLUG}.wav"ffmpeg -i "segments/${SLUG}.wav" -ac 1 -ar 16000 -c:a libopus -b:a 32k "segments/${SLUG}.ogg"curl -s https://api.groq.com/openai/v1/audio/transcriptions \
-H "Authorization: Bearer $GROQ_API_KEY" \
-F file="@segments/${SLUG}.ogg" \
-F model="whisper-large-v3-turbo" \
-F language="en" \
-F response_format="verbose_json" \
-F 'timestamp_granularities[]=segment'For archiving individual talk videos:
yt-dlp -f 91 \
--downloader ffmpeg \
--downloader-args "ffmpeg_i:-allowed_extensions ALL" \
--download-sections "*${START}-${END}" \
-o "clips/${SLUG}.%(ext)s" \
"$VIDEO_URL"Run with ThreadPoolExecutor(max_workers=2) -- more than 2 concurrent yt-dlp downloads tends to get throttled.
Send each raw transcript to Claude (or another LLM) for cleanup. Use 3 concurrent API calls.
System prompt for cleanup:
You are an expert transcript editor. Take this raw auto-caption transcript and produce a clean, readable document.
Rules:
1. KEEP all timestamps in [HH:MM:SS | +MM:SS] format. Include them every 30-60 seconds.
2. Fix transcription errors: proper nouns, technical terms, company names, jargon.
3. Add paragraph breaks at natural topic transitions.
4. Remove filler words (um, uh, like, you know) unless they add meaning.
5. Preserve the speaker's voice -- clean up, don't rewrite.
6. For multi-speaker segments, use [Speaker Name]: format.
Output format:
# Talk Title
**Speaker** -- Role/Company
**Event**: {event name and date}
## Key Points
- 4-8 bullet points of key takeaways
## Timestamped Reading Transcript
[Content with timestamps, paragraphs, and light line-wrapping for readability]Organize into:
transcripts/
raw/ -- unedited VTT-parsed or Whisper output
cleaned/ -- LLM-cleaned markdown
clips/ -- individual talk MP4s (optional)
talks.json -- manifest with metadata
reports/
talk-manifest.md -- summary tableFor the common case of a YouTube conference stream with chapters and auto-captions:
# 1. Grab metadata + captions
yt-dlp --write-info-json --skip-download --write-auto-sub --sub-lang en -o "media/%(id)s" "$URL"
# 2. Build talks.json from chapters in info.json
python3 scripts/build_transcripts.py
# 3. Download individual clips (parallel, optional)
python3 scripts/download_clips.py
# 4. Clean up transcripts with LLM
python3 scripts/cleanup_transcripts.pyThe build_transcripts.py and cleanup_transcripts.py scripts handle VTT parsing, deduplication, timestamp formatting, and LLM cleanup. See the AIE Europe project for reference implementations.
© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in conference-transcribe of swyxio/skills.
Open the folder on GitHubat commit 038ef34
Conference Transcribe next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Conference Transcribe this skillswyxio/skills | 176 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Native Subtitle Quote Imagechengyi-ai/native-subtitle-quote-image | 2.6k | — | ~2.4k | Automated safety check: Pass | MIT | |
| Video Dataoxylabs/agent-skills | 875 | — | ~1.4k | Automated safety check: Pass | MIT | |
| Summarizetrpc-group/trpc-agent-go | 1.9k | 22 repos | ~552 | Automated safety check: Pass | Apache-2.0 | |
| Youtube PublishAndonywang123/Epost | 197 | — | ~3.4k | Automated safety check: Warn | None | |
| Youtube Transcribe Skillfeiskyer/codex-settings | 244 | — | ~745 | Automated safety check: Pass | MIT |
chengyi-ai/native-subtitle-quote-image
将本地视频或用户有权处理的在线视频,经过来源获取、文字稿定位、选题选句、精确取帧、紧凑裁切、拼图和逐张质检,制作成 3:4 或保留画面原比例的视频字幕长图。支持两种明确分开的输出:保留画面内已烧录字幕的原生字幕模式,以及把已审核的时间点与台词绘制到真实视频帧上的脚本字幕模式。用户要求原生字幕截图、字幕帧拼图、YouTube…
oxylabs/agent-skills
YouTube data extraction API and high-bandwidth proxy downloads.
trpc-group/trpc-agent-go
Summarize or extract text/transcripts from URLs, podcasts, and local files (great fallback for “transcribe this YouTube/video”).
Andonywang123/Epost
Prepare an English YouTube release with local Chinese-to-English translation, subtitles and cover localization, then use a deterministic script connected to dedicated Chrome and YouTube Studio to…
feiskyer/codex-settings
Extract subtitles or a transcript from a YouTube URL and save normalized timestamped text locally.
oxbshw/watch-skill
The user shared a video URL, a YouTube/TikTok/stream link, a local video file, a screen recording, a meeting recording, or a playlist/folder of videos — "watch this", "summarize this video", "what's…
swyxio/skills
Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.
swyxio/skills
Design, implement, audit, or refresh protected username and handle namespaces for public products.
swyxio/skills
Fully automated new Mac setup for fullstack web developers and AI engineers.
swyxio/skills
Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…
swyxio/skills
Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.
swyxio/skills
Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.
Works with
Categories
Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts. Conference Transcribe is an agent skill from swyxio/skills. Transcribe a multi-talk conference livestream or long YouTube video into separate per-talk transcripts.
Conference Transcribe fits situations like: user says transcribe this conference; split this livestream into talks; transcribe each talk separately; provides a YouTube URL of a multi-hour event stream with chapter timestamps.
Run `npx skills add swyxio/skills --skill conference-transcribe -a claude-code`. Or copy the skill folder (conference-transcribe in swyxio/skills) into .claude/skills/conference-transcribe in your project. Claude Code loads it when a task matches its description.
Run `npx skills add swyxio/skills --skill conference-transcribe -a codex`. Or copy the skill folder (conference-transcribe in swyxio/skills) into .agents/skills/conference-transcribe in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill conference-transcribe -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/conference-transcribe, .gemini/skills/conference-transcribe, .github/skills/conference-transcribe and .opencode/skills/conference-transcribe in your project.
Going by SKILL.md and its folder, Conference Transcribe needs the command-line tools its instructions call (yt-dlp, ffmpeg, python3, uv, curl and pip) and credentials named HF_TOKEN, GROQ_API_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; A credential in GROQ_API_KEY; A credential in ANTHROPIC_API_KEY. Compatibility (from SKILL.md): Requires macOS with ffmpeg and yt-dlp installed. Needs at least one transcription backend (Groq API recommended for speed). Needs an LLM API key (Anthropic recommended) for cleanup pass. .
SKILL.md names 2 domains. In commands or code: youtube.com and api.groq.com; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Conference Transcribe is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Conference Transcribe: Native Subtitle Quote Image (chengyi-ai/native-subtitle-quote-image, 2.6k stars), Video Data (oxylabs/agent-skills, 875 stars), Summarize (trpc-group/trpc-agent-go, 1.9k stars) and Youtube Publish (Andonywang123/Epost, 197 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
swyxio (a GitHub user) maintains it in swyxio/skills, which has 176 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.
Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.