Agent skill

Baoyu Youtube Transcript

by JimLiu in JimLiu/baoyu-skills

Downloads YouTube video transcripts/subtitles and cover images by URL or video ID.

MITAuto-check passedMedia & Creative

Install Baoyu Youtube Transcript

skills CLI
$ npx skills add JimLiu/baoyu-skills --skill baoyu-youtube-transcript -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install JimLiu/baoyu-skills baoyu-youtube-transcript --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/JimLiu/baoyu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/baoyu-youtube-transcript .claude/skills/baoyu-youtube-transcript && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
baoyu-youtube-transcript
GitHub stars
26k
Used in
1 other repo
Token cost
~2.4k tokens
SKILL.md length
874 words
Files
9 (incl. scripts)
Skills in repo
22
Repo updated
First seen
Licence
MIT

At a glance

Downloads YouTube video transcripts/subtitles and cover images by URL or video ID.

  • Works in 5 steps: Run with --list first if the user hasn't… → Always single-quote the URL when running… → Default: run with --chapters --speakers… → …
  • User asks to get YouTube transcript
  • SKILL.md covers Script Directory, Usage, Options and Optional Environment Variables, plus 7 more sections
  • Runs TypeScript scripts from its folder; calls npx and yt-dlp; reaches youtube.com and youtu.be

What it does

Baoyu Youtube Transcript is an agent skill from JimLiu/baoyu-skills. Downloads YouTube video transcripts/subtitles and cover images by URL or video ID. Supports multiple languages, translation, chapters, and speaker identification. Caches raw data for fast re-formatting. Use when user asks to "get YouTube transcript", "download subtitles", "get captions", "YouTube字幕", "YouTube封面", "视频封面", "video thumbnail", "video cover image", or provides a YouTube URL and wants the transcript/subtitle text or cover image extracted.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `prompts/speaker-transcript.md`, `scripts/main.test.ts` and `scripts/main.ts`).

It sits in Media & Creative, covering Transcription, Video and podcast notes and Social media graphics. It works with YouTube. The licence is MIT.

When your agent uses it

  • User asks to get YouTube transcript
  • Download subtitles
  • Video thumbnail
  • Video cover image

Example prompts

  • “get YouTube transcript”
  • “download subtitles”
  • “get captions”
  • “/baoyu-youtube-transcript”

Requirements

  • Node.js

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Run with --list first if the user hasn't specified a language, to show available options
  2. Always single-quote the URL when running the script — zsh treats ? as a glob wildcard, so an unquoted YouTube URL causes "no matches…
  3. Default: run with --chapters --speakers for the richest output (chapters + speaker identification)
  4. The script auto-saves cached data + output file and prints the file path
  5. For --speakers mode: after the script saves the raw file, follow the speaker identification workflow below to post-process with speaker…

What it can do on your machine

Read from SKILL.md and the folder at commit 1567581. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 7 files in scripts/ (TypeScript), which the agent can run.

    Shell commands in SKILL.md call:

    • npx
    • yt-dlp

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com
    • youtu.be

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Baoyu Youtube Transcript loads about 2.4k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 874 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from JimLiu/baoyu-skills at commit 1567581, republished under its MIT licence (© JimLiu). 874 words, ~2,398 tokens.

Download SKILL.mdSave it as .claude/skills/baoyu-youtube-transcript/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
baoyu-youtube-transcript
description
Downloads YouTube video transcripts/subtitles and cover images by URL or video ID. Supports multiple languages, translation, chapters, and speaker identification. Caches raw data for fast re-formatting. Use when user asks to "get YouTube transcript", "download subtitles", "get captions", "YouTube字幕", "YouTube封面", "视频封面", "video thumbnail", "video cover image", or provides a YouTube URL and wants the transcript/subtitle text or cover image extracted.
version
1.1.0

YouTube Transcript

Downloads transcripts (subtitles/captions) from YouTube videos. Works with both manually created and auto-generated transcripts. No API key or browser required — uses YouTube's InnerTube API directly and automatically falls back to yt-dlp when YouTube blocks the direct API path.

Fetches video metadata and cover image on first run, caches raw data for fast re-formatting.

Script Directory

Scripts in scripts/ subdirectory. {baseDir} = this SKILL.md's directory path. Resolve ${BUN_X} runtime: if bun installed → bun; if npx available → npx -y bun; else suggest installing bun. Replace {baseDir} and ${BUN_X} with actual values.

ScriptPurpose
scripts/main.tsTranscript download CLI

Usage

bash
# Default: markdown with timestamps (English)
${BUN_X} {baseDir}/scripts/main.ts <youtube-url-or-id>

# Specify languages (priority order)
${BUN_X} {baseDir}/scripts/main.ts <url> --languages zh,en,ja

# Without timestamps
${BUN_X} {baseDir}/scripts/main.ts <url> --no-timestamps

# With chapter segmentation
${BUN_X} {baseDir}/scripts/main.ts <url> --chapters

# With speaker identification (requires AI post-processing)
${BUN_X} {baseDir}/scripts/main.ts <url> --speakers

# SRT subtitle file
${BUN_X} {baseDir}/scripts/main.ts <url> --format srt

# Translate transcript
${BUN_X} {baseDir}/scripts/main.ts <url> --translate zh-Hans

# List available transcripts
${BUN_X} {baseDir}/scripts/main.ts <url> --list

# Force re-fetch (ignore cache)
${BUN_X} {baseDir}/scripts/main.ts <url> --refresh

Options

OptionDescriptionDefault
<url-or-id>YouTube URL or video ID (multiple allowed)Required
--languages <codes>Language codes, comma-separated, in priority orderen
--format <fmt>Output format: text, srttext
--translate <code>Translate to specified language code
--listList available transcripts instead of fetching
--timestampsInclude [HH:MM:SS → HH:MM:SS] timestamps per paragraphon
--no-timestampsDisable timestamps
--chaptersChapter segmentation from video description
--speakersRaw transcript with metadata for speaker identification
--exclude-generatedSkip auto-generated transcripts
--exclude-manually-createdSkip manually created transcripts
--refreshForce re-fetch, ignore cached data
-o, --output <path>Save to specific file pathauto-generated
--output-dir <dir>Base output directoryyoutube-transcript

Optional Environment Variables

VariableDescription
YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSERPassed to yt-dlp --cookies-from-browser during fallback, e.g. chrome, safari, firefox, or chrome:Profile 1

Input Formats

Accepts any of these as video input:

  • Full URL: https://www.youtube.com/watch?v=dQw4w9WgXcQ
  • Short URL: https://youtu.be/dQw4w9WgXcQ
  • Embed URL: https://www.youtube.com/embed/dQw4w9WgXcQ
  • Shorts URL: https://www.youtube.com/shorts/dQw4w9WgXcQ
  • Video ID: dQw4w9WgXcQ

Output Formats

FormatExtensionDescription
text.mdMarkdown with frontmatter (incl. description), title heading, summary, optional TOC/cover/timestamps/chapters/speakers
srt.srtSubRip subtitle format for video players

Output Directory

youtube-transcript/
├── .index.json                          # Video ID → directory path mapping (for cache lookup)
└── {channel-slug}/{title-full-slug}/
    ├── meta.json                        # Video metadata (title, channel, description, duration, chapters, etc.)
    ├── transcript-raw.json              # Raw transcript snippets from YouTube API (cached)
    ├── transcript-sentences.json        # Sentence-segmented transcript (split by punctuation, merged across snippets)
    ├── imgs/
    │   └── cover.jpg                    # Video thumbnail
    ├── transcript.md                    # Markdown transcript (generated from sentences)
    └── transcript.srt                   # SRT subtitle (generated from raw snippets, if --format srt)
  • {channel-slug}: Channel name in kebab-case
  • {title-full-slug}: Full video title in kebab-case

The --list mode outputs to stdout only (no file saved).

Caching

On first fetch, the script saves:

  • meta.json — video metadata, chapters, cover image path, language info
  • transcript-raw.json — raw transcript snippets from YouTube API ({ text, start, duration }[])
  • transcript-sentences.json — sentence-segmented transcript ({ text, start: "HH:mm:ss", end: "HH:mm:ss" }[]), split by sentence-ending punctuation (.?!…。?! etc.), timestamps proportionally allocated by character length, CJK-aware text merging
  • imgs/cover.jpg — video thumbnail

Subsequent runs for the same video use cached data (no network calls). Use --refresh to force re-fetch. If a different language is requested, the cache is automatically refreshed.

When YouTube returns anti-bot / blocked responses on the direct InnerTube path, the script retries with alternate client identities and then falls back to yt-dlp if available. If fallback is needed but yt-dlp is unavailable, the agent should decide how to make yt-dlp available and continue rather than pushing the installation decision to the user.

SRT output (--format srt) is generated from transcript-raw.json. Text/markdown output uses transcript-sentences.json for natural sentence boundaries.

Workflow

When user provides a YouTube URL and wants the transcript:

  1. Run with --list first if the user hasn't specified a language, to show available options
  2. Always single-quote the URL when running the script — zsh treats ? as a glob wildcard, so an unquoted YouTube URL causes "no matches found": use 'https://www.youtube.com/watch?v=ID'
  3. Default: run with --chapters --speakers for the richest output (chapters + speaker identification)
  4. The script auto-saves cached data + output file and prints the file path
  5. For --speakers mode: after the script saves the raw file, follow the speaker identification workflow below to post-process with speaker labels

When user only wants a cover image or metadata, running the script with any option will also cache meta.json and imgs/cover.jpg.

When re-formatting the same video (e.g., first text then SRT), the cached data is reused — no re-fetch needed.

Show full SKILL.md (299 more words)Show less

Chapter & Speaker Workflow

Chapters (--chapters)

The script parses chapter timestamps from the video description (e.g., 0:00 Introduction), segments the transcript by chapter boundaries, groups snippets into readable paragraphs, and saves as .md with a Table of Contents. No further processing needed.

If no chapter timestamps exist in the description, the transcript is output as grouped paragraphs without chapter headings.

Speaker Identification (--speakers)

Speaker identification requires AI processing. The script outputs a raw .md file containing:

  • YAML frontmatter with video metadata (title, channel, date, cover, description, language)
  • Video description (for speaker name extraction)
  • Chapter list from description (if available)
  • Raw transcript in SRT format (pre-computed start/end timestamps, token-efficient)

After the script saves the raw file, spawn a sub-agent (use a cheaper model like Sonnet for cost efficiency) to process speaker identification:

  1. Read the saved .md file
  2. Read the prompt template at {baseDir}/prompts/speaker-transcript.md
  3. Process the raw transcript following the prompt:
    • Identify speakers using video metadata (title → guest, channel → host, description → names)
    • Detect speaker turns from conversation flow, question-answer patterns, and contextual cues
    • Segment into chapters (use description chapters if available, else create from topic shifts)
    • Format with **Speaker Name:** labels, paragraph grouping (2-4 sentences), and [HH:MM:SS → HH:MM:SS] timestamps
  4. Overwrite the .md file with the processed transcript (keep the YAML frontmatter)

When --speakers is used, --chapters is implied — the processed output always includes chapter segmentation.

Error Cases

ErrorMeaning
Transcripts disabledVideo has no captions at all
No transcript foundRequested language not available
Video unavailableVideo deleted, private, or region-locked
IP blockedToo many requests, try again later
Age restrictedVideo requires login for age verification
bot detectedThe script retries alternate clients and then yt-dlp; if fallback tooling is missing, the agent should resolve that itself, otherwise if it still fails try YOUTUBE_TRANSCRIPT_COOKIES_FROM_BROWSER=safari (or your browser)

© JimLiu, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts) in skills/baoyu-youtube-transcript of JimLiu/baoyu-skills.

  • SKILL.md
  • prompts/speaker-transcript.md
  • scripts/main.test.ts
  • scripts/main.ts
  • scripts/shared.ts
  • scripts/storage.ts
  • scripts/transcript.ts
  • scripts/types.ts
  • scripts/youtube.ts

Open the folder on GitHubat commit 1567581

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in JimLiu/baoyu-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Baoyu Youtube Transcript next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Baoyu Youtube Transcript compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Baoyu Youtube Transcript this skillJimLiu/baoyu-skills26k1 repos~2.4kAutomated safety check: PassMIT
Video To NotesKIRVO-REPORTING/video-to-notes105—~1.5kAutomated safety check: PassMIT
Video Link Transcript ExtractorSpaceZephyr/creator-buddy1.6k—~498Automated safety check: PassNone
Watch Videocoreyhaines31/makerskills850—~3.8kAutomated safety check: PassMIT
Video Transcript Downloadersundial-org/awesome-openclaw-skills6632 repos~574Automated safety check: PassNone
Video SummaryLeoYeAI/openclaw-master-skills2.2k—~4.2kAutomated safety check: PassMIT

Similar skills

  • Video To Notes

    KIRVO-REPORTING/video-to-notes

    Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.

    105 GitHub stars~1.5k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Video Link Transcript Extractor

    SpaceZephyr/creator-buddy

    Extracts subtitles or a full transcript from YouTube, Xiaoyuzhou, Bilibili, Douyin and Xiaohongshu links, using speech recognition when a video has no subtitles.

    1.6k GitHub stars~498 tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    850 GitHub stars~3.8k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Video Transcript Downloader

    sundial-org/awesome-openclaw-skills

    Download videos, audio, subtitles, and clean paragraph-style transcripts from YouTube and any other yt-dlp supported site.

    663 GitHub starsUsed in 2 repos~574 tokens
    Media & CreativeAuto-check passed
  • Video Summary

    LeoYeAI/openclaw-master-skills

    Video summarization for Bilibili, Xiaohongshu, Douyin, and YouTube.

    2.2k GitHub stars~4.2k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Gemini Yt Video Transcript

    sundial-org/awesome-openclaw-skills

    Create a verbatim transcript for a YouTube URL using Google Gemini (speaker labels, paragraph breaks; no time codes).

    663 GitHub stars~293 tokensUpdated 7 mo ago
    Media & CreativeAuto-check passed

More from JimLiu/baoyu-skills

All 22 skills in this repo
  • Markdown Article Formatter

    JimLiu/baoyu-skills

    Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.

    26k GitHub starsUsed in 6 repos~3.5k tokens
    Auto-check passed
  • X to Markdown Converter

    JimLiu/baoyu-skills

    Saves tweets, threads and X Articles as Markdown files with YAML front matter, using an unofficial API that asks for your consent first.

    26k GitHub starsUsed in 4 repos~1.8k tokens
    Auto-check: warnings
  • SVG Diagram Generator

    JimLiu/baoyu-skills

    Creates standalone dark-themed SVG diagrams, including architecture, flowchart, sequence, structural, mind map, timeline and state machine types.

    26k GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Publishes articles and image-text posts to a WeChat Official Account through the API or Chrome CDP, converting markdown to WeChat-ready HTML with link citations.

    26k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check: warnings
  • Image Compressor

    JimLiu/baoyu-skills

    Compresses images to WebP by default, or to PNG or JPEG, picking the best available tool on the machine and optionally processing whole folders.

    26k GitHub starsUsed in 5 repos~598 tokens
    Auto-check passed
  • Generates text and images through an unofficial, reverse-engineered Gemini Web API, supporting reference images and multi-turn conversations.

    26k GitHub starsUsed in 5 repos~1.6k tokens
    Auto-check passed

Works with

Questions about Baoyu Youtube Transcript

What does Baoyu Youtube Transcript do?

Downloads YouTube video transcripts/subtitles and cover images by URL or video ID. Baoyu Youtube Transcript is an agent skill from JimLiu/baoyu-skills. Downloads YouTube video transcripts/subtitles and cover images by URL or video ID.

When should I use Baoyu Youtube Transcript?

Baoyu Youtube Transcript fits situations like: user asks to get YouTube transcript; download subtitles; video thumbnail; video cover image.

How do I install Baoyu Youtube Transcript in Claude Code?

Run `npx skills add JimLiu/baoyu-skills --skill baoyu-youtube-transcript -a claude-code`. Or copy the skill folder (skills/baoyu-youtube-transcript in JimLiu/baoyu-skills) into .claude/skills/baoyu-youtube-transcript in your project. Claude Code loads it when a task matches its description.

How do I install Baoyu Youtube Transcript in Codex?

Run `npx skills add JimLiu/baoyu-skills --skill baoyu-youtube-transcript -a codex`. Or copy the skill folder (skills/baoyu-youtube-transcript in JimLiu/baoyu-skills) into .agents/skills/baoyu-youtube-transcript in your project. Codex loads it when a task matches its description.

Can I use Baoyu Youtube Transcript in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add JimLiu/baoyu-skills --skill baoyu-youtube-transcript -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/baoyu-youtube-transcript, .gemini/skills/baoyu-youtube-transcript, .github/skills/baoyu-youtube-transcript and .opencode/skills/baoyu-youtube-transcript in your project.

What does Baoyu Youtube Transcript need to run?

Going by SKILL.md and its folder, Baoyu Youtube Transcript needs TypeScript for the scripts in its folder and the command-line tools its instructions call (npx and yt-dlp). Our summary lists: Node.js.

Does Baoyu Youtube Transcript access the network?

SKILL.md names 2 domains. In commands or code: youtube.com and youtu.be; the agent is likely to contact these when it follows the instructions. This is read from the text; nothing was executed.

Is Baoyu Youtube Transcript safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Baoyu Youtube Transcript use?

Baoyu Youtube Transcript is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Baoyu Youtube Transcript use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Baoyu Youtube Transcript?

Skills that share tags, products or a category with Baoyu Youtube Transcript: Video To Notes (KIRVO-REPORTING/video-to-notes, 105 stars), Video Link Transcript Extractor (SpaceZephyr/creator-buddy, 1.6k stars), Watch Video (coreyhaines31/makerskills, 850 stars) and Video Transcript Downloader (sundial-org/awesome-openclaw-skills, 663 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Baoyu Youtube Transcript?

JimLiu (a GitHub user) maintains it in JimLiu/baoyu-skills, which has 26,455 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on September 10, 2026.

Source: JimLiu/baoyu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.