Agent skill

YouTube Talk Notetaker

by dair-ai in dair-ai/dair-academy-plugins

Converts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server.

MITAuto-check passedKnowledge Management

Install YouTube Talk Notetaker

skills CLI
$ npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install dair-ai/dair-academy-plugins youtube-notetaker --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/youtube-notetaker/skills/youtube-notetaker .claude/skills/youtube-notetaker && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
youtube-notetaker
GitHub stars
614
Token cost
~2.3k tokens
SKILL.md length
919 words
Files
11 (incl. scripts)
Skills in repo
8
Repo updated
First seen
Licence
MIT

At a glance

Converts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server.

  • Works in 8 steps: Resolve the id and check embeddability → Download video + subtitles → Detect candidate slide timestamps → …
  • Studying a conference talk and capturing its slides with timestamps
  • SKILL.md covers Architecture (read this first), Requirements, Adding a video — the pipeline and Markdown file shape (reference), plus 2 more sections
  • Runs Python and Shell scripts from its folder; calls python3 and pip; reaches youtube.com

What it does

Each YouTube video you add becomes one plain markdown file in a local library folder, which defaults to ~/video-deepdives/ and can be changed with VIDEO_LIBRARY_DIR. The file holds video metadata and a slides list in its frontmatter and the full transcript as timestamped lines in the body, while slide images are saved in a media folder named per video. No database or cloud service is involved, and the markdown files stay the single source of truth.

Helper scripts download the video, detect and extract slides at their timestamps, build a contact sheet, convert subtitles to a transcript and write the library item. A bundled Python server, serve.py, renders the library as a front page index plus a per-video view with the slide deck, embedded player and searchable transcript, and notes edited there are saved back to the markdown files. It needs yt-dlp, ffmpeg, Pillow and PyYAML.

When your agent uses it

  • Studying a conference talk and capturing its slides with timestamps
  • Building a personal library of YouTube talks as markdown files
  • Taking searchable, timestamped notes on a video

Example prompts

  • “Add this YouTube video to my study library and extract the slides at their timestamps.”
  • “Take timestamped notes on this conference talk and save them in my video library.”
  • “Start the library server on port 8000 so I can browse my saved talks.”

Requirements

  • yt-dlp and ffmpeg on the PATH
  • Python 3 with Pillow and PyYAML

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Resolve the id and check embeddability
  2. Download video + subtitles
  3. Detect candidate slide timestamps
  4. Build a contact sheet and CURATE
  5. Extract the curated slides at full quality and install to _media
  6. Build the transcript
  7. Write notes and assemble the markdown file
  8. Serve and verify (always do this)

What it can do on your machine

Read from SKILL.md and the folder at commit 0abffdc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

YouTube Talk Notetaker loads about 2.3k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 919 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~224
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from dair-ai/dair-academy-plugins at commit 0abffdc, republished under its MIT licence (© dair-ai). 919 words, ~2,303 tokens.

Download SKILL.mdSave it as .claude/skills/youtube-notetaker/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
youtube-notetaker
description
Turn any YouTube talk into a studyable, interactive deep-dive stored as plain markdown: extract slide images at their timestamps, build a clean [HH:MM:SS] transcript, write editable notes, and save everything as one markdown file per video in a local library folder. A bundled zero-dependency server (serve.py) renders the whole library as a single-page artifact (front-page index + per-video split pane: slide deck + embedded player + searchable transcript) with notes that save back to the markdown files. Fully self-contained: no external services, configurable library path. Use when the user gives a YouTube URL and wants to study a talk, capture slides, take timestamped notes, or build a talk library. Triggers on: "deep-dive this talk", "extract slides from this video", "add this YouTube video to my study library", "take notes on this talk", followed by a YouTube URL.

YouTube Notetaker

Build a personal library of YouTube talks you study with. Each video becomes one plain markdown file: slide snapshots at their timestamps, a full timestamped transcript, and editable notes. A small bundled server renders the library as an interactive deep-dive in the browser. No database, no cloud service. Everything is files on disk you fully own.

Architecture (read this first)

The markdown library is the single source of truth. The artifact is a thin HTML shell that fetches from the server and writes notes back. Never hardcode video data into the HTML.

  • Library: a plain folder, set by VIDEO_LIBRARY_DIR (default ~/video-deepdives/).
    • One markdown file per video, filename slug = YouTube id (e.g. RtywqDFBYnQ.md).
    • Frontmatter holds video metadata + a slides array.
    • Body holds the full transcript as [HH:MM:SS] text lines.
    • _media/ holds slide images, namespaced per video as <youtube_id>-slide-NN.jpg to avoid collisions between videos.
  • Server: scripts/serve.py, a single stdlib + PyYAML file. Start it with:
    python3 scripts/serve.py --dir ~/video-deepdives --port 8000
    It serves the artifact at / and a small API the artifact talks to:
    • GET /api/video-deepdives (front page fetches this) lists every video.
    • GET /api/video-deepdives/<id> returns one video {meta, body}.
    • GET /api/video-deepdives/_media/<file> serves a slide image.
    • PATCH /api/video-deepdives/<id> with {fields:{slides:[...]}} writes notes back.
    • It picks up new videos automatically the moment a markdown file exists. Adding a video means writing a markdown file + media; you almost never touch the HTML.
    • The /api/video-deepdives URL namespace is local to the bundled server.
  • Artifact: reference/artifact.html, served by serve.py at /. A clean reference copy; only rewrite it if the user wants a UI change. For new videos, leave it alone.

Requirements

  • yt-dlp and ffmpeg on PATH (download + frame/scene extraction).
  • Python 3 with Pillow (contact sheet) and PyYAML (markdown file + server).
    pip install yt-dlp pillow pyyaml      # ffmpeg via your package manager

Adding a video — the pipeline

All helper scripts are in scripts/. Work in a scratch dir (e.g. /tmp/ytnote-<id>/), then copy final assets into the library. Set VIDEO_LIBRARY_DIR once per shell if you don't want the default. Do not use em dashes (—) or arrows (→) in notes/titles.

1. Resolve the id and check embeddability
scripts/setup.sh "<youtube_url_or_id>"

Prints the 11-char YTID, the scratch dir, the target library path, and whether YouTube embedding is allowed (oembed 200) or blocked (oembed 401, e.g. some university talks). If blocked, inline playback won't work but the artifact degrades gracefully to an "open at this moment on YouTube" link, so proceed normally.

2. Download video + subtitles
scripts/download.sh "<YTID>" /tmp/ytnote-<YTID>

Uses yt-dlp to grab the video (≤720p is plenty for slide frames) and the best available subtitles (manual if present, else auto-captions) as .vtt. Also fetches title/uploader.

3. Detect candidate slide timestamps
scripts/detect_slides.sh /tmp/ytnote-<YTID>/video.mp4 /tmp/ytnote-<YTID>

Runs ffmpeg scene detection (select='gt(scene,0.3)') and writes scene_times.txt (seconds). 0.3 is a good default; lower it (0.2) for subtle slide decks, raise it (0.4) for busy video.

4. Build a contact sheet and CURATE
python3 scripts/contact_sheet.py /tmp/ytnote-<YTID>/video.mp4 /tmp/ytnote-<YTID>/scene_times.txt /tmp/ytnote-<YTID>/contact.jpg

Read contact.jpg (labeled with index + timestamp). This is the human-judgment step: keep frames that are real content slides; drop talking-head shots, transitions, duplicates, and blurry mid-animation frames. Save the kept timestamps (seconds) to /tmp/ytnote-<YTID>/keep.txt, one per line. Typical talk yields 15-25 slides.

Show full SKILL.md (426 more words)Show less
5. Extract the curated slides at full quality and install to _media
python3 scripts/extract_slides.py <YTID> /tmp/ytnote-<YTID>/video.mp4 /tmp/ytnote-<YTID>/keep.txt > /tmp/ytnote-<YTID>/slides.json

Extracts each kept timestamp at 1280px wide, JPEG, and copies them into $VIDEO_LIBRARY_DIR/_media/ as <YTID>-slide-01.jpg, -02.jpg, … (numbered in time order). Progress goes to stderr; a clean slides.json scaffold prints to stdout, so redirect it to a file as shown, then fill in title and note.

Tip: talks are often a slide + speaker-cam composite, and speakers flip back and forth, so the same slide appears at several timestamps. Keep the cleanest instance of each, and re-anchor each slide's t to where it is actually discussed in the transcript (better "play from here" UX).

6. Build the transcript
python3 scripts/vtt_to_transcript.py /tmp/ytnote-<YTID>/*.vtt /tmp/ytnote-<YTID>/transcript.txt

Parses the VTT into clean, de-duplicated [HH:MM:SS] text lines (YouTube auto-captions repeat rolling text; the script collapses it). This becomes the markdown body.

7. Write notes and assemble the markdown file

For each kept slide, write a 1-3 sentence note grounded in the transcript around that timestamp (don't invent claims). Then assemble:

python3 scripts/write_library_item.py \
  --id <YTID> \
  --title "Talk title" \
  --speaker "Name, Role, Org" \
  --tags tag1,tag2,tag3 \
  --slides /tmp/ytnote-<YTID>/slides.json \
  --transcript /tmp/ytnote-<YTID>/transcript.txt

Writes $VIDEO_LIBRARY_DIR/<YTID>.md with correct frontmatter + body.

8. Serve and verify (always do this)
python3 scripts/serve.py --dir "$VIDEO_LIBRARY_DIR" --port 8000 &
scripts/verify.sh <YTID>                 # defaults to http://127.0.0.1:8000

verify.sh curls the collection list, the item, the first slide image, and the artifact, asserting HTTP 200 and that the new id appears in the index. Then open http://127.0.0.1:8000/#/<YTID> in a browser to confirm slides + transcript + notes render.

Markdown file shape (reference)

markdown
---
id: RtywqDFBYnQ
title: Memory and dreaming for self-learning agents
youtube_id: RtywqDFBYnQ
speaker: Mahesh, Product Manager, Platform team at Anthropic
source_url: https://www.youtube.com/watch?v=RtywqDFBYnQ
slide_count: 19
created: '2026-05-25'
tags: [anthropic, memory, agents]
slides:
- idx: 1
  t: 55.7                 # seconds (float ok), used for seeking
  mmss: 00:55             # display label
  title: Agent primitives have evolved
  note: One to three sentences grounded in the transcript at this timestamp.
  img: /api/video-deepdives/_media/RtywqDFBYnQ-slide-01.jpg
# ... more slides
---
## Transcript
[00:00:08] Hello, everyone...
[00:00:11] ...

Notes:

  • idx can be sparse/non-contiguous; the artifact sorts slides by t, so ordering is by timestamp, not idx.
  • img is always a /api/video-deepdives/_media/<file> URL (served by serve.py), never base64.
  • Slide note is what the user edits in the UI; PATCH writes the whole slides array back.

Gotchas

  • Embedding disabled (oembed 401): inline player is blocked by the video owner. Not a bug; the artifact shows an "open at this moment on YouTube" link instead. Mention it to the user.
  • Image collisions: always namespace media <YTID>-slide-NN.jpg. Never reuse bare slide-NN.jpg for a new video.
  • Auto-caption noise: rolling YouTube captions duplicate text across cues; use the provided VTT parser, don't dump raw VTT into the body.
  • Don't touch existing videos when adding a new one. Each video is an independent file.
  • Server not picking up a video: confirm the .md file is directly inside --dir (not a subfolder) and the filename is <YTID>.md.

What makes this portable

  • No orchestrator / no database. Storage is a plain folder of markdown + images.
  • One env var (VIDEO_LIBRARY_DIR) controls where the library lives.
  • One small server file (serve.py, stdlib + PyYAML) renders everything and handles note write-back. Drop it anywhere Python runs.
  • The markdown files are portable: readable in Obsidian or any editor, and the frontmatter is standard YAML.

© dair-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (scripts) in plugins/youtube-notetaker/skills/youtube-notetaker of dair-ai/dair-academy-plugins.

  • SKILL.md
  • reference/artifact.html
  • scripts/contact_sheet.py
  • scripts/detect_slides.sh
  • scripts/download.sh
  • scripts/extract_slides.py
  • scripts/serve.py
  • scripts/setup.sh
  • scripts/verify.sh
  • scripts/vtt_to_transcript.py
  • scripts/write_library_item.py

Open the folder on GitHubat commit 0abffdc

Compare with similar skills

YouTube Talk Notetaker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

YouTube Talk Notetaker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
YouTube Talk Notetaker this skilldair-ai/dair-academy-plugins614—~2.3kAutomated safety check: PassMIT
Video Lenskar2phi/video-lens113—~8.2kAutomated safety check: NotesMIT
Youtube Transcriptbrowser-act/skills6.1k—~2.1kAutomated safety check: PassMIT
URL to Markdown FetcherJimLiu/baoyu-skills27k1 repos~2.1kAutomated safety check: PassMIT
Video To Notelike-attract/video-to-note124—~1kAutomated safety check: PassMIT
Gbro Series Vocabpyang5166/gbro-series-vocab163—~1.1kAutomated safety check: PassMIT

Similar skills

  • Video Lens

    kar2phi/video-lens

    Fetch a YouTube transcript and generate an executive summary, key points, and timestamped topic list as a polished HTML report.

    113 GitHub stars~8.2k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check: notes
  • Youtube Transcript

    browser-act/skills

    YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…

    6.1k GitHub stars~2.1k tokensUpdated 1 mo ago
    Knowledge ManagementAuto-check passed
  • URL to Markdown Fetcher

    JimLiu/baoyu-skills

    Fetches a web page, X post, YouTube transcript or Hacker News thread through a Chrome-driven CLI and saves it as clean markdown.

    27k GitHub starsUsed in 1 repo~2.1k tokens
    Knowledge ManagementAuto-check passed
  • Video To Note

    like-attract/video-to-note

    Generate structured, timestamped Markdown notes from videos (Bilibili, Douyin, YouTube, or local media files) using the local VideoToNo service.

    124 GitHub stars~1k tokensUpdated 7 days ago
    Knowledge ManagementAuto-check passed
  • Gbro Series Vocab

    pyang5166/gbro-series-vocab

    追剧学英语 / Learn English vocabulary from TV series. An agent skill from pyang5166/gbro-series-vocab.

    163 GitHub stars~1.1k tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Video To Notes

    KIRVO-REPORTING/video-to-notes

    Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.

    105 GitHub stars~1.5k tokensUpdated 1 mo ago
    Media & CreativeAuto-check passed

More from dair-ai/dair-academy-plugins

All 8 skills in this repo
  • Research Wiki Builder

    dair-ai/dair-academy-plugins

    Creates and maintains configurable research wikis: scaffold a folder, add sources, compile pages and indexes, and file query answers back.

    614 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Gemini Image Generator

    dair-ai/dair-academy-plugins

    Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.

    614 GitHub starsUsed in 2 repos~3.5k tokens
    Auto-check: notes
  • Survey Paper Generator

    dair-ai/dair-academy-plugins

    Builds a single-file HTML survey paper on an AI or ML topic from a research bundle the agent curates, with prose and SVG figures written by Kimi K2.6.

    614 GitHub starsUsed in 2 repos~2.1k tokens
    Auto-check: notes
  • Learn

    dair-ai/dair-academy-plugins

    Help a user learn a topic through adaptive tutoring, lesson planning, practice, retrieval checks, explanations, study guides, or exercises.

    614 GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check passed
  • LLM Council on Fireworks AI

    dair-ai/dair-academy-plugins

    Has several open-weight models answer a question, rank each other's anonymized answers, then lets a chairman model write the final response through Fireworks AI.

    614 GitHub stars~5k tokensUpdated 2 mo ago
    Auto-check: notes
  • X Agent Intelligence Feed

    dair-ai/dair-academy-plugins

    Builds a self-contained HTML digest of AI and agent news pulled from chosen X accounts through the official X MCP server, grouped into categories like Coding Agents and Agent Research.

    614 GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed

Questions about YouTube Talk Notetaker

What does YouTube Talk Notetaker do?

Converts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server. Each YouTube video you add becomes one plain markdown file in a local library folder, which defaults to ~/video-deepdives/ and can be changed with VIDEO_LIBRARY_DIR. The file holds video metadata and a slides list in its frontmatter and the full transcript as timestamped lines in the body, while slide images are saved in a media folder named per video.

When should I use YouTube Talk Notetaker?

YouTube Talk Notetaker fits situations like: studying a conference talk and capturing its slides with timestamps; building a personal library of YouTube talks as markdown files; taking searchable, timestamped notes on a video.

How do I install YouTube Talk Notetaker in Claude Code?

Run `npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a claude-code`. Or copy the skill folder (plugins/youtube-notetaker/skills/youtube-notetaker in dair-ai/dair-academy-plugins) into .claude/skills/youtube-notetaker in your project. Claude Code loads it when a task matches its description.

How do I install YouTube Talk Notetaker in Codex?

Run `npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a codex`. Or copy the skill folder (plugins/youtube-notetaker/skills/youtube-notetaker in dair-ai/dair-academy-plugins) into .agents/skills/youtube-notetaker in your project. Codex loads it when a task matches its description.

Can I use YouTube Talk Notetaker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/youtube-notetaker, .gemini/skills/youtube-notetaker, .github/skills/youtube-notetaker and .opencode/skills/youtube-notetaker in your project.

What does YouTube Talk Notetaker need to run?

Going by SKILL.md and its folder, YouTube Talk Notetaker needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: yt-dlp and ffmpeg on the PATH; Python 3 with Pillow and PyYAML.

Does YouTube Talk Notetaker access the network?

SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is YouTube Talk Notetaker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does YouTube Talk Notetaker use?

YouTube Talk Notetaker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does YouTube Talk Notetaker use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to YouTube Talk Notetaker?

Skills that share tags, products or a category with YouTube Talk Notetaker: Video Lens (kar2phi/video-lens, 113 stars), Youtube Transcript (browser-act/skills, 6.1k stars), URL to Markdown Fetcher (JimLiu/baoyu-skills, 27k stars) and Video To Note (like-attract/video-to-note, 124 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains YouTube Talk Notetaker?

dair-ai (a GitHub organization) maintains it in dair-ai/dair-academy-plugins, which has 614 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on July 21, 2026.

Source: dair-ai/dair-academy-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.