Agent skill

Youtube Notes

by nicknisi in nicknisi/claude-plugins

Pull captions from a YouTube video and turn them into chapter-aligned notes where every claim carries a clickable timestamp.

MITAuto-check passedWriting & Content

Install Youtube Notes

skills CLI
$ npx skills add nicknisi/claude-plugins --skill youtube-notes -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nicknisi/claude-plugins youtube-notes --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nicknisi/claude-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/sources/skills/youtube-notes .claude/skills/youtube-notes && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
youtube-notes
GitHub stars
114
Token cost
~2.8k tokens
SKILL.md length
1,278 words
Files
6 (incl. scripts)
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Pull captions from a YouTube video and turn them into chapter-aligned notes where every claim carries a clickable timestamp.

  • Works in 3 steps: Pull --mode triage to see the chapter… → For a short or medium video, pull --mode… → For something multi-hour, pull the…
  • Wants to summarize
  • SKILL.md covers Fetch, The transcript is cached, so…, The main use: answering… and Mode by intent, plus 9 more sections
  • Runs TypeScript and Shell scripts from its folder; calls node and brew; reaches youtube.com

What it does

Youtube Notes is an agent skill from nicknisi/claude-plugins. Pull captions from a YouTube video and turn them into chapter-aligned notes where every claim carries a clickable timestamp. Use this whenever a youtube.com or youtu.be link appears, including a bare pasted link with no instructions attached, and whenever the user wants to summarize, digest, triage, or take notes on a video, talk, lecture, keynote, podcast, or interview: "is this worth watching", "TL;DR this talk", "what do they say about X", "find the part where they discuss Y", "save this to my notes", "get me…

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts (for example `scripts/fetch_video.ts`, `test/fixtures/captions.json` and `test/fixtures/metadata.json`).

It sits in Writing & Content, covering Summarization. It works with YouTube. The repository describes itself as: Nick's own marketplace of Claude Plugins. The licence is MIT.

When your agent uses it

  • Wants to summarize
  • Take notes on a video
  • Interview: is this worth watching
  • What do they say about X

Example prompts

  • “is this worth watching”
  • “TL;DR this talk”
  • “what do they say about X”
  • “/youtube-notes”

Requirements

  • Node.js
  • A Bash shell

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Pull --mode triage to see the chapter map and size.
  2. For a short or medium video, pull --mode full and keep it in context. Every
  3. For something multi-hour, pull the chapters that bear on the question with

What it can do on your machine

Read from SKILL.md and the folder at commit 6a6decd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (TypeScript and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • node
    • brew

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Youtube Notes loads about 2.8k tokens when it runs. Until then it costs about 183 tokens; SKILL.md has 1,278 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~183
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from nicknisi/claude-plugins at commit 6a6decd, republished under its MIT licence (© nicknisi). 1,278 words, ~2,752 tokens.

Download SKILL.mdSave it as .claude/skills/youtube-notes/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
youtube-notes
description
Pull captions from a YouTube video and turn them into chapter-aligned notes where every claim carries a clickable timestamp. Use this whenever a youtube.com or youtu.be link appears, including a bare pasted link with no instructions attached, and whenever the user wants to summarize, digest, triage, or take notes on a video, talk, lecture, keynote, podcast, or interview: "is this worth watching", "TL;DR this talk", "what do they say about X", "find the part where they discuss Y", "save this to my notes", "get me the transcript". Prefer this over generic web fetching or page readers for YouTube links, because the page HTML carries no transcript. Not for downloading video or audio files, and not for local media.

YouTube Notes

Turn a video into something citable. The bundled script does the mechanical work — fetching captions, aligning them to the uploader's chapters, merging caption slivers into readable blocks — so your job is synthesis and judgment.

Fetch

One TypeScript file, run directly by Node with no build step. The --help output is the authoritative flag reference:

bash
node ${CLAUDE_PLUGIN_ROOT}/skills/youtube-notes/scripts/fetch_video.ts --help

It shells out to uvx for youtube-transcript-api (captions) and yt-dlp (metadata, chapters). Both run on demand; nothing is installed globally. If uvx is missing the script exits 2 and tells the user to brew install uv.

The transcript is cached, so re-running is free

The first fetch for a video is cached by id. Every later call for that video reads from disk and touches no network: 35s becomes 0.06s, and switching modes or pulling different chapters costs nothing.

This shapes how to work. Do not ration calls to avoid refetching, because after the first one there is no refetch. Pull triage to see the shape, then pull the chapters you actually need, then pull different ones when the next question lands. The only real budget is context, not network.

It is also the main defense against rate limiting, which is the failure mode this skill hits most. A request you never make cannot be throttled.

--refresh forces a refetch, which you need only if a video's captions were genuinely republished. --no-cache skips it entirely.

The main use: answering questions about a video

Most requests here are conversational. Someone wants to interrogate a video, not receive a document. Expect several questions across a conversation about one video.

Load once, properly, then answer from what you have:

  1. Pull --mode triage to see the chapter map and size.
  2. For a short or medium video, pull --mode full and keep it in context. Every follow-up is then answered without another call.
  3. For something multi-hour, pull the chapters that bear on the question with --chapters, and pull more as the conversation moves. Cached, so cheap.

Answer with the deep links inline so the user can jump to the source. When the video does not address something, say so plainly rather than reaching for an adjacent passage.

Resist turning every question into a digest. If someone asks what the speaker thinks about X, answer that, cite it, and stop.

Mode by intent

User intentMode
"is this worth watching", "what's this about", bare pasted linktriage
"summarize", "take notes", "digest this", "save to my notes"full
a question about the content, or a conversation about the videotriage, then full (or --chapters <n>)
"give me the transcript"transcript
bash
node ${CLAUDE_PLUGIN_ROOT}/skills/youtube-notes/scripts/fetch_video.ts "<url>" --mode triage

Long videos

Chapter indices are stable across calls, so an index from triage is safe to reuse later. A 3h45m podcast is ~75k tokens in full and ~2k in triage, so target chapters rather than loading everything:

bash
node ${CLAUDE_PLUGIN_ROOT}/skills/youtube-notes/scripts/fetch_video.ts "<url>" \
  --mode full --chapters 3,7 --out /tmp/video.json

The script prints an estimated token count to stderr and warns when output is large. Above roughly 25k tokens, write to --out and read the file back rather than piping it through stdout.

When captions fail

Two failures, one remedy:

  • Exit 4, rate limited. YouTube throttles the caption endpoint per IP. It is transient and affects only captions; metadata, chapters, and the audio stream keep working.
  • Exit 3, no captions published. Permanent for that video.

Both are answered by --whisper-fallback, which downloads the audio and transcribes it locally. It works during a block because the media CDN is not subject to the caption quota. Verified: youtube-transcript-api, yt-dlp subtitles, and curl with a real browser User-Agent all get 429 on the caption endpoint while audio downloads at full speed.

Ask before using it on the first video of a session. It costs a one-time model download of roughly 1.6 GB and a minute or two of compute, which is a real cost the user should agree to. Once they have agreed, keep using it for that conversation without asking again.

bash
node ${CLAUDE_PLUGIN_ROOT}/skills/youtube-notes/scripts/fetch_video.ts "<url>" \
  --mode full --whisper-fallback

Do not suggest browser cookies as a workaround. Upstream warns that authenticating that way eventually gets the account permanently banned.

Whisper output is cached like any other transcript, so the cost is paid once per video.

Show full SKILL.md (594 more words)Show less

What comes back

JSON modes return metadata plus a chapters array. Fields that drive decisions:

  • chapters_source — youtube (the uploader's own chapters, trustworthy structure), time-sliced (no chapters, so 5-minute buckets titled by range), or single (short video, one bucket). Only youtube reflects authorial intent; do not present time-sliced bucket titles as if they were chapter names.

  • caption_kind — where the words came from, which sets how much to trust an exact quote:

    • manual — a human-authored track. Quote freely.
    • generated — YouTube's ASR. Mishears proper nouns, technical terms, and numbers.
    • whisper — local ASR, not anything YouTube served. Usually cleaner prose than generated, but still guesses at names, product names, and version numbers.

    For generated or whisper, say once that quotes are approximate, and flag specific proper nouns you suspect rather than silently correcting them into something plausible.

  • word_count and reading_minutes — compare reading_minutes against duration_hms to tell the user whether reading beats watching. It usually does for interviews and usually doesn't for anything visual.

  • chapters_total vs chapters_included — if these differ you are looking at a subset, so do not claim to have covered the whole video.

  • chapters[].blocks[] — each block is { t, s, text } where t is a display timestamp and s is that timestamp in whole seconds.

Citation rule

Every claim you attribute to the video gets a deep link built from the block's s value:

<url>&t=<s>

So a claim from a block with "s": 754 on video abc12345678 cites as https://www.youtube.com/watch?v=abc12345678&t=754.

Never invent a timestamp. If you cannot locate a claim in a specific block, attribute it to the chapter and use the chapter's link, or leave it uncited.

Triage output

Short. The user is deciding whether to spend 40 minutes.

markdown
**[Title]** — Channel · 18:40 · ~3,400 words

[Two sentences: what it actually covers, and how it treats the subject.]

**Worth it if:** [who this serves]
**Skip if:** [who it wastes]
**Best chapters:** [Chapter name](link) · [Chapter name](link)

Give an actual verdict. "It depends on your goals" is not a verdict.

Digest output

Ordered so someone who stops reading after ten seconds still got the most valuable part.

markdown
# [Title]

**Channel:** [name] · **Duration:** [hms] · **Published:** [date]
**Source:** [url]

## TL;DR

[One or two sentences carrying the video's actual thesis.]

## Key takeaways

- [3-7 standalone insights. Each should survive being read alone.]

## Claims worth checking

- [Claim] — [[m:ss]](deep link)
- [Mark anything the speaker asserts without support, and anything you know to
  be contested.]

## Walkthrough

### [Chapter title] ([m:ss](link))

[What happens here, in a few sentences. Follow the video's own chapters when
`chapters_source` is `youtube`.]

## Notable quotes

> "[quote]" — [[m:ss]](deep link)

## Open questions

[What the video raises and does not answer.]

Skip sections that would be filler. A ten-minute tutorial does not need "Claims worth checking."

Extract mode

When the user asks one question, answer only that question. Do not emit a digest. Triage first to read the chapter titles, then pull just the plausible chapters with --chapters. Reply with the answer, the deep links backing it, and an explicit "the video does not address this" when it doesn't.

Saving to notes

Only write files when the user asks. If they do, look for a vault directory in .claude/sources.local.md:

markdown
---
vault_dir: /path/to/vault/Videos
---

Absent that setting, ask once where notes should live, then offer to record the answer there.

When writing into a vault, match the conventions already in it — read a neighboring note before inventing frontmatter. A reasonable default:

markdown
---
tags:
  - video
  - youtube
youtube_id: '<video id>'
source: '<url>'
channel: '<channel>'
date: <YYYY-MM-DD>
---

Quote youtube_id so it stays a string, and grep for it before writing to avoid duplicating a video already captured.

Honesty

  • The transcript is the only evidence you have. You did not watch the video, so do not describe visuals, slides, or demos beyond what the captions state.
  • Auto-generated captions garble names and numbers. Do not silently "fix" a term into something plausible; flag the uncertainty instead.
  • Never search the web for a third-party transcript. A re-publication of unknown provenance cannot be verified, and --whisper-fallback gets you the real audio anyway.
  • Do not retry a rate-limited fetch in a loop. It extends the block. Offer --whisper-fallback or a wait.
  • Metadata and chapters survive a caption block, so a title-and-chapter-map answer is available even when the transcript is not. Say clearly that it comes from metadata and that you have not read the content.
  • When caption_kind is whisper, the words are a local transcription of the audio. Do not present them as YouTube's captions.

© nicknisi, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (scripts) in plugins/sources/skills/youtube-notes of nicknisi/claude-plugins.

  • SKILL.md
  • scripts/fetch_video.ts
  • test/fixtures/captions.json
  • test/fixtures/metadata.json
  • test/run.sh
  • test/uvx

Open the folder on GitHubat commit 6a6decd

Compare with similar skills

Youtube Notes next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Youtube Notes compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Youtube Notes this skillnicknisi/claude-plugins114—~2.8kAutomated safety check: PassMIT
Summarize Anythingswyxio/skills176—~6.3kAutomated safety check: PassMIT
YouTube Transcript ReformatterPrismer-AI/PrismerCloud1.6k2 repos~905Automated safety check: PassMIT
Video Lenskar2phi/video-lens113—~8.2kAutomated safety check: NotesMIT
YouTube Transcripts and SearchZeroPointRepo/youtube-skills1k1 repos~3.1kAutomated safety check: PassMIT
Youtube Transcript Extractor API Skillbrowser-act/skills6.1k1 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Summarize Anything

    swyxio/skills

    Summarizes arbitrarily long text (1k-1M words) using recursive map-reduce with any LLM backend.

    176 GitHub stars~6.3k tokensUpdated 5 days ago
    Writing & ContentAuto-check passed
  • YouTube Transcript Reformatter

    Prismer-AI/PrismerCloud

    Fetches a YouTube transcript with a helper script and reshapes it into chapters, summaries, X threads, blog posts or timestamped quotes.

    1.6k GitHub starsUsed in 2 repos~905 tokens
    Knowledge ManagementAuto-check passed
  • Video Lens

    kar2phi/video-lens

    Fetch a YouTube transcript and generate an executive summary, key points, and timestamped topic list as a polished HTML report.

    113 GitHub stars~8.2k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check: notes
  • YouTube Transcripts and Search

    ZeroPointRepo/youtube-skills

    Fetches YouTube transcripts and searches videos, channels and playlists through the TranscriptAPI.com service.

    1k GitHub starsUsed in 1 repo~3.1k tokens
    Knowledge ManagementAuto-check passed
  • This skill helps users automatically extract YouTube video transcripts and metadata via the BrowserAct API.

    6.1k GitHub starsUsed in 1 repo~1.2k tokens
    Knowledge ManagementAuto-check passed
  • News Aggregator Skill

    cclank/news-aggregator-skill

    Comprehensive news aggregator that fetches, filters, and deeply analyzes real-time content from 44+ sources including Hacker News, Lobsters, Dev.to, GitHub, arXiv, Hugging Face Papers, AIHOT, TLDR…

    1.3k GitHub stars~2.1k tokensUpdated 4 mo ago
    Writing & ContentAuto-check passed

More from nicknisi/claude-plugins

All 12 skills in this repo
  • Blog Post Writer

    nicknisi/claude-plugins

    Turn Nick Nisi's raw material into a blog post draft in his voice.

    114 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Image Gen

    nicknisi/claude-plugins

    Generate or edit images via Google Gemini (nano-banana-pro) or OpenAI gpt-image-2.

    114 GitHub stars~502 tokensUpdated 2 mo ago
    Auto-check passed
  • Prototype

    nicknisi/claude-plugins

    Build a throwaway prototype to answer a design question before committing to real implementation.

    114 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Squad Review

    nicknisi/claude-plugins

    Review the current branch with six specialist lenses (security, correctness, conventions, tests, architecture, duplication), then put every finding through an adversarial verifier that tries to…

    114 GitHub stars~751 tokensUpdated 2 mo ago
    Auto-check passed
  • Tmux

    nicknisi/claude-plugins

    Read and drive other tmux panes when Claude runs inside tmux.

    114 GitHub stars~3.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Conference Talk Builder

    nicknisi/claude-plugins

    Create conference talk outlines and slide-by-slide content plans using narrative frameworks.

    114 GitHub stars~2.7k tokensUpdated 2 mo ago
    Auto-check passed

Works with

Questions about Youtube Notes

What does Youtube Notes do?

Pull captions from a YouTube video and turn them into chapter-aligned notes where every claim carries a clickable timestamp. Youtube Notes is an agent skill from nicknisi/claude-plugins. Pull captions from a YouTube video and turn them into chapter-aligned notes where every claim carries a clickable timestamp.

When should I use Youtube Notes?

Youtube Notes fits situations like: wants to summarize; take notes on a video; interview: is this worth watching; what do they say about X.

How do I install Youtube Notes in Claude Code?

Run `npx skills add nicknisi/claude-plugins --skill youtube-notes -a claude-code`. Or copy the skill folder (plugins/sources/skills/youtube-notes in nicknisi/claude-plugins) into .claude/skills/youtube-notes in your project. Claude Code loads it when a task matches its description.

How do I install Youtube Notes in Codex?

Run `npx skills add nicknisi/claude-plugins --skill youtube-notes -a codex`. Or copy the skill folder (plugins/sources/skills/youtube-notes in nicknisi/claude-plugins) into .agents/skills/youtube-notes in your project. Codex loads it when a task matches its description.

Can I use Youtube Notes in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nicknisi/claude-plugins --skill youtube-notes -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/youtube-notes, .gemini/skills/youtube-notes, .github/skills/youtube-notes and .opencode/skills/youtube-notes in your project.

What does Youtube Notes need to run?

Going by SKILL.md and its folder, Youtube Notes needs TypeScript and a shell for the scripts in its folder and the command-line tools its instructions call (node and brew). Our summary lists: Node.js; A Bash shell.

Does Youtube Notes access the network?

SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Youtube Notes safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Youtube Notes use?

Youtube Notes is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Youtube Notes use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Youtube Notes?

Skills that share tags, products or a category with Youtube Notes: Summarize Anything (swyxio/skills, 176 stars), YouTube Transcript Reformatter (Prismer-AI/PrismerCloud, 1.6k stars), Video Lens (kar2phi/video-lens, 113 stars) and YouTube Transcripts and Search (ZeroPointRepo/youtube-skills, 1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Youtube Notes?

nicknisi (a GitHub user) maintains it in nicknisi/claude-plugins, which has 114 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on August 11, 2026.

Source: nicknisi/claude-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.