Agent skill

Youtube Transcript

by browser-act in browser-act/skills

YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…

MITAuto-check passedKnowledge Management

Install Youtube Transcript

skills CLI
$ npx skills add browser-act/skills --skill youtube-transcript -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install browser-act/skills youtube-transcript --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/browser-act/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/solutions/video-platforms/youtube-transcript .claude/skills/youtube-transcript && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
youtube-transcript
GitHub stars
6.1k
Token cost
~2.1k tokens
SKILL.md length
845 words
Files
4 (incl. scripts)
Skills in repo
87
Repo updated
First seen
Licence
MIT

At a glance

YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…

  • Works in 5 steps: navigate… → eval "$(python… → eval "$(python… → …
  • The user shares a YouTube URL
  • SKILL.md covers Language, Objective, Prerequisites and Pre-execution Checks, plus 7 more sections
  • Runs Python scripts from its folder; calls python; reaches youtube.com

What it does

Youtube Transcript is an agent skill from browser-act/skills. YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into summaries, chapter outlines, Twitter/X threads, blog posts, or notable quotes. Use when the user shares a YouTube URL or video link, asks to summarize a video, get a transcript, extract content from a YouTube video, get YouTube captions, extract YouTube captions, download YouTube captions, transcribe YouTube video, YouTube…

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `scripts/extract-transcript-segments.py`, `scripts/get-languages.py` and `scripts/open-transcript-panel.py`).

It sits in Knowledge Management, covering Video and podcast notes, Transcription and Blog and article writing. It works with YouTube, X (Twitter) and Python. The repository describes itself as: Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session… The licence is MIT.

When your agent uses it

  • The user shares a YouTube URL
  • Asks to summarize a video
  • Get a transcript
  • Extract content from a YouTube video

Example prompts

  • “/youtube-transcript”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. navigate https://www.youtube.com/watch?v={VIDEO_ID} → wait stable
  2. eval "$(python scripts/get-languages.py)" — confirm transcripts are available; note the language list
  3. eval "$(python scripts/open-transcript-panel.py)" — open the panel
  4. wait stable — wait for panel content to load
  5. eval "$(python scripts/extract-transcript-segments.py)" — extract all segments

What it can do on your machine

Read from SKILL.md and the folder at commit 11c057b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Youtube Transcript loads about 2.1k tokens when it runs. Until then it costs about 216 tokens; SKILL.md has 845 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~216
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from browser-act/skills at commit 11c057b, republished under its MIT licence (© browser-act). 845 words, ~2,140 tokens.

Download SKILL.mdSave it as .claude/skills/youtube-transcript/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
youtube-transcript
description
YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into summaries, chapter outlines, Twitter/X threads, blog posts, or notable quotes. Use when the user shares a YouTube URL or video link, asks to summarize a video, get a transcript, extract content from a YouTube video, get YouTube captions, extract YouTube captions, download YouTube captions, transcribe YouTube video, YouTube video to text, make a thread from YouTube, YouTube to blog post, YouTube to article, pull transcript from YouTube, YouTube content extraction, convert YouTube to text, video to transcript. Also applies when user wants to reformat any YouTube video content into structured output (chapters, threads, blog articles, key quotes).

YouTube — Transcript Extraction & Content Reformatting

YouTube video URL → timestamped transcript → summary / chapters / thread / blog / quotes

Language

All process output to user (progress updates, process notifications) follows the user's language.

Objective

Extract the full transcript from a YouTube video's built-in transcript panel, then transform it into the output format the user requests.

Prerequisites

  • Target YouTube video page is already open in the browser: https://www.youtube.com/watch?v={VIDEO_ID}

Pre-execution Checks

1. Tool Readiness

If browser-act has been confirmed available in the current session → skip this step.

Invoke browser-act via Skill tool to load usage. If installation or configuration issues arise, follow its guidance to resolve then retry.

Capability Components

This Skill's operational boundary = what the user can manually do in their browser. It only reads data already displayed to the user on the page, never bypassing authentication or access controls. JS code is encapsulated in Python files under the scripts/ directory, invoked via eval "$(python scripts/xxx.py)". Use the bash tool for execution.

DOM: Check transcript availability and list languages

eval "$(python scripts/get-languages.py)"

No parameters. Reads ytInitialPlayerResponse from the current page.

Output example:

json
{
  "available_languages": [
    {"code": "en", "name": "English", "kind": "manual", "is_auto": false},
    {"code": "en", "name": "English (auto-generated)", "kind": "asr", "is_auto": true}
  ],
  "count": 2
}

Returns {"error": true, "message": "..."} when transcripts are disabled or page is not a YouTube video.

DOM: Open transcript panel

eval "$(python scripts/open-transcript-panel.py)"

No parameters. Clicks the "Show transcript" button below the video (handles multiple UI language variants automatically for robustness).

Must call wait stable after this to allow the panel to fully load.

Output example:

json
{"success": true, "label": "内容转文字"}
DOM: Extract all transcript segments

eval "$(python scripts/extract-transcript-segments.py)"

No parameters. Scrolls the open transcript panel to trigger lazy loading for long videos, then extracts all segments.

Output example:

json
{
  "segment_count": 24,
  "segments": [
    {"ts": "0:18", "text": "We're no strangers to love"},
    {"ts": "0:27", "text": "You know the rules and so do I"}
  ],
  "full_text": "We're no strangers to love You know the rules...",
  "timestamped_text": "0:18 We're no strangers to love\n0:27 You know the rules..."
}
Composite: Full transcript fetch workflow
  1. navigate https://www.youtube.com/watch?v={VIDEO_ID} → wait stable
  2. eval "$(python scripts/get-languages.py)" — confirm transcripts are available; note the language list
  3. eval "$(python scripts/open-transcript-panel.py)" — open the panel
  4. wait stable — wait for panel content to load
  5. eval "$(python scripts/extract-transcript-segments.py)" — extract all segments

Use timestamped_text from the output as input for the Transform step below.

Transform: Content Reformatting

After fetching the transcript, transform it based on what the user requests. If the user did not specify a format, default to the Full Document — output all five sections in order.

  • Summary: Concise 5–10 sentence overview of the entire video
  • Chapters: Group by topic shifts, output timestamped chapter list
  • Thread: Twitter/X thread format — numbered posts, each under 280 characters
  • Blog post: Full article with title, H2 sections per major topic, key quotes, and takeaways
  • Quotes: Notable quotes with their timestamps

Default Full Document output order (when no specific format is requested):

  1. Summary
  2. Chapters
  3. Thread
  4. Blog Post
  5. Quotes
Workflow
  1. Fetch transcript using the Composite component above.
  2. Validate: confirm segment_count >= 1. If empty, tell the user the video has transcripts disabled.
  3. Chunk if needed: if full_text exceeds ~50,000 characters, split timestamped_text into overlapping chunks (~40K characters with 2K overlap) and summarize each chunk before merging.
  4. Transform into the requested format(s) using the timestamped_text field. If no format specified, produce all five sections.
  5. Verify: re-read the output for coherence, correct timestamps (if chapters), and completeness before presenting.
Example — Chapters Output
0:00 Introduction — host opens with the problem statement
3:45 Background — prior work and why existing solutions fall short
12:20 Core method — walkthrough of the proposed approach
24:10 Results — benchmark comparisons and key takeaways
31:55 Q&A — audience questions on scalability and next steps
Show full SKILL.md (338 more words)Show less
Example — Thread Output
1/ Just watched an incredible video on [topic]. Key takeaways 🧵

2/ First insight: [point]. This matters because [reason].

3/ The surprising part: [finding]. Most assume [belief], but this shows otherwise.

4/ Practical takeaway: [action].

5/ Full video: [URL]

Error Handling

  • Transcripts disabled: get-languages.py returns error; tell user and suggest checking if captions are available on the video page
  • Private/unavailable video: page will not load correctly; relay the error and ask user to verify the URL
  • Transcript button not found: usually means the user is not on a video page, or the page hasn't finished loading; navigate to the URL and retry
  • No segments after panel opens: retry open-transcript-panel.py + wait stable + extract-transcript-segments.py once

Known Limitations

  • Language selection: the transcript panel shows the language YouTube defaults to for the user's region. Switching to a specific language requires changing the caption language in the player's CC settings first; automatic language switching is not implemented.
  • Auto-generated transcripts (kind: asr) may have lower accuracy than manual captions.
  • Videos that require login to view will not have a transcript panel accessible.

Execution Efficiency

  • Batch orchestration: Write a bash script to loop through video URLs serially within a single session — navigate to each video, run the 3-step composite workflow, save result, then move to the next. Do not parallelize within one browser. To increase throughput for large batches, open multiple stealth browser sessions and distribute URLs across them.
  • Test before batch execution: After writing a batch script, first test with 1–2 videos to confirm the full workflow runs correctly; only then run the full batch.
  • Reduce redundant pre-operations: Pre-execution checks (tool readiness) only need to run once per session; skip them for subsequent videos in the same batch.
  • Error resumption: Save each video's result immediately after extraction; on failure, resume from the failed video rather than starting over.

Success Criteria

segment_count >= 1 AND full_text length > 0

Experience Notes

Path: {working-directory}/browser-act-skill-forge-memories/youtube-content-youtube-transcript.memory.md

Before execution: If the file exists, read it first — it records unexpected situations encountered during past executions (e.g., a strategy has become ineffective); adjust strategy order accordingly.

After execution: If an unexpected situation is encountered (strategy became ineffective, page redesigned, anti-scraping upgraded, better path discovered), append a line: {YYYY-MM-DD}: {what happened} → {conclusion}

Normal execution does not write to the file.

© browser-act, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in solutions/video-platforms/youtube-transcript of browser-act/skills.

  • SKILL.md
  • scripts/extract-transcript-segments.py
  • scripts/get-languages.py
  • scripts/open-transcript-panel.py

Open the folder on GitHubat commit 11c057b

Compare with similar skills

Youtube Transcript next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Youtube Transcript compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Youtube Transcript this skillbrowser-act/skills6.1k—~2.1kAutomated safety check: PassMIT
Video Lenskar2phi/video-lens112—~8.2kAutomated safety check: NotesMIT
Video Analysisericosiu/ai-marketing-skills3.6k—~1.3kAutomated safety check: PassMIT
Youtube FetcherJimmySadek/youtube-fetcher-to-markdown485—~1.8kAutomated safety check: PassMIT
Youtube Transcriptintellectronica/agent-skills2952 repos~394Automated safety check: PassCC0-1.0
Social Media Managementmanojbajaj95/claude-gtm-plugin1041 repos~3.9kAutomated safety check: PassMIT

Similar skills

  • Video Lens

    kar2phi/video-lens

    Fetch a YouTube transcript and generate an executive summary, key points, and timestamped topic list as a polished HTML report.

    112 GitHub stars~8.2k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check: notes
  • Video Analysis

    ericosiu/ai-marketing-skills

    Analyze YouTube videos or local footage using transcripts first, with silent video checks for demonstrations, delivery, editing, and clip boundaries.

    3.6k GitHub stars~1.3k tokensUpdated 14 days ago
    Knowledge ManagementAuto-check passed
  • Youtube Fetcher

    JimmySadek/youtube-fetcher-to-markdown

    Retrieve YouTube transcripts and subtitles, summarize or analyze what was said, or save an Obsidian-ready Markdown knowledge-base note with captions, creator metadata, chapters, language, and source…

    485 GitHub stars~1.8k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • Youtube Transcript

    intellectronica/agent-skills

    Extract transcripts from YouTube videos. An agent skill from intellectronica/agent-skills.

    295 GitHub starsUsed in 2 repos~394 tokens
    Knowledge ManagementAuto-check passed
  • Social Media Management

    manojbajaj95/claude-gtm-plugin

    Comprehensive social media management for all platforms (LinkedIn, Twitter/X, Instagram, TikTok, Facebook, Pinterest, YouTube).

    104 GitHub starsUsed in 1 repo~3.9k tokens
    Writing & ContentAuto-check passed
  • Repurpose

    Blotato-Inc/blotato-skills

    Turns one long-form input (blog post, newsletter, YouTube transcript, or script) into 3 platform-native LinkedIn posts, 5 X/Twitter threads, and 2 short-form video scripts for Reels/TikTok.

    182 GitHub stars~1.7k tokensUpdated 1 mo ago
    Writing & ContentAuto-check passed

More from browser-act/skills

All 87 skills in this repo
  • Amazon ASIN Lookup

    browser-act/skills

    Fetches structured Amazon product details such as title, price, ratings and availability for a given ASIN through BrowserAct's lookup API template.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Amazon Best Sellers Finder

    browser-act/skills

    Extracts structured Amazon product data for a keyword and marketplace through the BrowserAct API, including titles, prices, ratings, reviews, sales volume and promotions.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Amazon Buy Box Monitor

    browser-act/skills

    Pulls Amazon product details, competing seller prices and seller ratings for a given ASIN through the BrowserAct API, without browser automation.

    6.1k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed
  • Analyzes a competitor's Amazon listing by ASIN with BrowserAct data extraction, then reports what it does well, where the market has gaps and opportunity points for your own listing.

    6.1k GitHub starsUsed in 2 repos~3.2k tokens
    Auto-check passed
  • Pulls structured Amazon search results (titles, ASINs, prices, ratings, specifications) for a keyword and brand through BrowserAct's Amazon Product API template.

    6.1k GitHub starsUsed in 2 repos~1.5k tokens
    Auto-check passed
  • Collects structured product data from Amazon search results for a keyword and optional brand, using a BrowserAct script, for market and competitor research.

    6.1k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check passed

Questions about Youtube Transcript

What does Youtube Transcript do?

YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…. Youtube Transcript is an agent skill from browser-act/skills. YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into summaries, chapter outlines, Twitter/X threads, blog posts, or notable quotes.

When should I use Youtube Transcript?

Youtube Transcript fits situations like: the user shares a YouTube URL; asks to summarize a video; get a transcript; extract content from a YouTube video.

How do I install Youtube Transcript in Claude Code?

Run `npx skills add browser-act/skills --skill youtube-transcript -a claude-code`. Or copy the skill folder (solutions/video-platforms/youtube-transcript in browser-act/skills) into .claude/skills/youtube-transcript in your project. Claude Code loads it when a task matches its description.

How do I install Youtube Transcript in Codex?

Run `npx skills add browser-act/skills --skill youtube-transcript -a codex`. Or copy the skill folder (solutions/video-platforms/youtube-transcript in browser-act/skills) into .agents/skills/youtube-transcript in your project. Codex loads it when a task matches its description.

Can I use Youtube Transcript in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add browser-act/skills --skill youtube-transcript -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/youtube-transcript, .gemini/skills/youtube-transcript, .github/skills/youtube-transcript and .opencode/skills/youtube-transcript in your project.

What does Youtube Transcript need to run?

Going by SKILL.md and its folder, Youtube Transcript needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Youtube Transcript access the network?

SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Youtube Transcript safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Youtube Transcript use?

Youtube Transcript is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Youtube Transcript use?

About 2.1k tokens (SKILL.md is roughly 8.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Youtube Transcript?

Skills that share tags, products or a category with Youtube Transcript: Video Lens (kar2phi/video-lens, 112 stars), Video Analysis (ericosiu/ai-marketing-skills, 3.6k stars), Youtube Fetcher (JimmySadek/youtube-fetcher-to-markdown, 485 stars) and Youtube Transcript (intellectronica/agent-skills, 295 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Youtube Transcript?

browser-act (a GitHub organization) maintains it in browser-act/skills, which has 6,108 GitHub stars. The repository holds 87 skills in this directory. The repository was last updated on August 24, 2026.

Source: browser-act/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.