Video Lens
kar2phi/video-lens
Fetch a YouTube transcript and generate an executive summary, key points, and timestamped topic list as a polished HTML report.
Converts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server.
$ npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install dair-ai/dair-academy-plugins youtube-notetaker --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/youtube-notetaker/skills/youtube-notetaker .claude/skills/youtube-notetaker && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "youtube-notetaker" agent skill from https://github.com/dair-ai/dair-academy-plugins/tree/main/plugins/youtube-notetaker/skills/youtube-notetaker into .claude/skills/youtube-notetaker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "youtube-notetaker", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/dair-ai/dair-academy-plugins/tree/main/plugins/youtube-notetaker/skills/youtube-notetakerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install dair-ai/dair-academy-plugins youtube-notetaker --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/youtube-notetaker/skills/youtube-notetaker .agents/skills/youtube-notetaker && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "youtube-notetaker" agent skill from https://github.com/dair-ai/dair-academy-plugins/tree/main/plugins/youtube-notetaker/skills/youtube-notetaker into .agents/skills/youtube-notetaker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "youtube-notetaker", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install dair-ai/dair-academy-plugins youtube-notetaker --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/youtube-notetaker/skills/youtube-notetaker .cursor/skills/youtube-notetaker && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "youtube-notetaker" agent skill from https://github.com/dair-ai/dair-academy-plugins/tree/main/plugins/youtube-notetaker/skills/youtube-notetaker into .cursor/skills/youtube-notetaker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "youtube-notetaker", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/dair-ai/dair-academy-plugins.git --path plugins/youtube-notetaker/skills/youtube-notetaker--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install dair-ai/dair-academy-plugins youtube-notetaker --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/youtube-notetaker/skills/youtube-notetaker .gemini/skills/youtube-notetaker && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "youtube-notetaker" agent skill from https://github.com/dair-ai/dair-academy-plugins/tree/main/plugins/youtube-notetaker/skills/youtube-notetaker into .gemini/skills/youtube-notetaker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "youtube-notetaker", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install dair-ai/dair-academy-plugins youtube-notetakerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/youtube-notetaker/skills/youtube-notetaker .github/skills/youtube-notetaker && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "youtube-notetaker" agent skill from https://github.com/dair-ai/dair-academy-plugins/tree/main/plugins/youtube-notetaker/skills/youtube-notetaker into .github/skills/youtube-notetaker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "youtube-notetaker", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install dair-ai/dair-academy-plugins youtube-notetaker --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/dair-ai/dair-academy-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/youtube-notetaker/skills/youtube-notetaker .opencode/skills/youtube-notetaker && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "youtube-notetaker" agent skill from https://github.com/dair-ai/dair-academy-plugins/tree/main/plugins/youtube-notetaker/skills/youtube-notetaker into .opencode/skills/youtube-notetaker/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "youtube-notetaker", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
youtube-notetakerConverts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server.
Each YouTube video you add becomes one plain markdown file in a local library folder, which defaults to ~/video-deepdives/ and can be changed with VIDEO_LIBRARY_DIR. The file holds video metadata and a slides list in its frontmatter and the full transcript as timestamped lines in the body, while slide images are saved in a media folder named per video. No database or cloud service is involved, and the markdown files stay the single source of truth.
Helper scripts download the video, detect and extract slides at their timestamps, build a contact sheet, convert subtitles to a transcript and write the library item. A bundled Python server, serve.py, renders the library as a front page index plus a per-video view with the slide deck, embedded player and searchable transcript, and notes edited there are saved back to the markdown files. It needs yt-dlp, ffmpeg, Pillow and PyYAML.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 0abffdc. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 9 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
python3pipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
youtube.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
YouTube Talk Notetaker loads about 2.3k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 919 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from dair-ai/dair-academy-plugins at commit 0abffdc, republished under its MIT licence (© dair-ai). 919 words, ~2,303 tokens.
.claude/skills/youtube-notetaker/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.Build a personal library of YouTube talks you study with. Each video becomes one plain markdown file: slide snapshots at their timestamps, a full timestamped transcript, and editable notes. A small bundled server renders the library as an interactive deep-dive in the browser. No database, no cloud service. Everything is files on disk you fully own.
The markdown library is the single source of truth. The artifact is a thin HTML shell that fetches from the server and writes notes back. Never hardcode video data into the HTML.
VIDEO_LIBRARY_DIR (default ~/video-deepdives/).RtywqDFBYnQ.md).slides array.[HH:MM:SS] text lines._media/ holds slide images, namespaced per video as <youtube_id>-slide-NN.jpg
to avoid collisions between videos.scripts/serve.py, a single stdlib + PyYAML file. Start it with:python3 scripts/serve.py --dir ~/video-deepdives --port 8000/ and a small API the artifact talks to:GET /api/video-deepdives (front page fetches this) lists every video.GET /api/video-deepdives/<id> returns one video {meta, body}.GET /api/video-deepdives/_media/<file> serves a slide image.PATCH /api/video-deepdives/<id> with {fields:{slides:[...]}} writes notes back./api/video-deepdives URL namespace is local to the bundled server.reference/artifact.html, served by serve.py at /. A clean reference copy;
only rewrite it if the user wants a UI change. For new videos, leave it alone.yt-dlp and ffmpeg on PATH (download + frame/scene extraction).Pillow (contact sheet) and PyYAML (markdown file + server).pip install yt-dlp pillow pyyaml # ffmpeg via your package managerAll helper scripts are in scripts/. Work in a scratch dir (e.g. /tmp/ytnote-<id>/), then
copy final assets into the library. Set VIDEO_LIBRARY_DIR once per shell if you don't want the
default. Do not use em dashes (—) or arrows (→) in notes/titles.
scripts/setup.sh "<youtube_url_or_id>"Prints the 11-char YTID, the scratch dir, the target library path, and whether YouTube
embedding is allowed (oembed 200) or blocked (oembed 401, e.g. some university talks).
If blocked, inline playback won't work but the artifact degrades gracefully to an "open at this
moment on YouTube" link, so proceed normally.
scripts/download.sh "<YTID>" /tmp/ytnote-<YTID>Uses yt-dlp to grab the video (≤720p is plenty for slide frames) and the best available
subtitles (manual if present, else auto-captions) as .vtt. Also fetches title/uploader.
scripts/detect_slides.sh /tmp/ytnote-<YTID>/video.mp4 /tmp/ytnote-<YTID>Runs ffmpeg scene detection (select='gt(scene,0.3)') and writes scene_times.txt (seconds).
0.3 is a good default; lower it (0.2) for subtle slide decks, raise it (0.4) for busy video.
python3 scripts/contact_sheet.py /tmp/ytnote-<YTID>/video.mp4 /tmp/ytnote-<YTID>/scene_times.txt /tmp/ytnote-<YTID>/contact.jpgRead contact.jpg (labeled with index + timestamp). This is the human-judgment step: keep
frames that are real content slides; drop talking-head shots, transitions, duplicates, and
blurry mid-animation frames. Save the kept timestamps (seconds) to /tmp/ytnote-<YTID>/keep.txt,
one per line. Typical talk yields 15-25 slides.
python3 scripts/extract_slides.py <YTID> /tmp/ytnote-<YTID>/video.mp4 /tmp/ytnote-<YTID>/keep.txt > /tmp/ytnote-<YTID>/slides.jsonExtracts each kept timestamp at 1280px wide, JPEG, and copies them into
$VIDEO_LIBRARY_DIR/_media/ as <YTID>-slide-01.jpg, -02.jpg, … (numbered in time order).
Progress goes to stderr; a clean slides.json scaffold prints to stdout, so redirect it to a
file as shown, then fill in title and note.
Tip: talks are often a slide + speaker-cam composite, and speakers flip back and forth, so the
same slide appears at several timestamps. Keep the cleanest instance of each, and re-anchor each
slide's t to where it is actually discussed in the transcript (better "play from here" UX).
python3 scripts/vtt_to_transcript.py /tmp/ytnote-<YTID>/*.vtt /tmp/ytnote-<YTID>/transcript.txtParses the VTT into clean, de-duplicated [HH:MM:SS] text lines (YouTube auto-captions repeat
rolling text; the script collapses it). This becomes the markdown body.
For each kept slide, write a 1-3 sentence note grounded in the transcript around that timestamp
(don't invent claims). Then assemble:
python3 scripts/write_library_item.py \
--id <YTID> \
--title "Talk title" \
--speaker "Name, Role, Org" \
--tags tag1,tag2,tag3 \
--slides /tmp/ytnote-<YTID>/slides.json \
--transcript /tmp/ytnote-<YTID>/transcript.txtWrites $VIDEO_LIBRARY_DIR/<YTID>.md with correct frontmatter + body.
python3 scripts/serve.py --dir "$VIDEO_LIBRARY_DIR" --port 8000 &
scripts/verify.sh <YTID> # defaults to http://127.0.0.1:8000verify.sh curls the collection list, the item, the first slide image, and the artifact,
asserting HTTP 200 and that the new id appears in the index. Then open
http://127.0.0.1:8000/#/<YTID> in a browser to confirm slides + transcript + notes render.
---
id: RtywqDFBYnQ
title: Memory and dreaming for self-learning agents
youtube_id: RtywqDFBYnQ
speaker: Mahesh, Product Manager, Platform team at Anthropic
source_url: https://www.youtube.com/watch?v=RtywqDFBYnQ
slide_count: 19
created: '2026-05-25'
tags: [anthropic, memory, agents]
slides:
- idx: 1
t: 55.7 # seconds (float ok), used for seeking
mmss: 00:55 # display label
title: Agent primitives have evolved
note: One to three sentences grounded in the transcript at this timestamp.
img: /api/video-deepdives/_media/RtywqDFBYnQ-slide-01.jpg
# ... more slides
---
## Transcript
[00:00:08] Hello, everyone...
[00:00:11] ...Notes:
idx can be sparse/non-contiguous; the artifact sorts slides by t, so ordering is by
timestamp, not idx.img is always a /api/video-deepdives/_media/<file> URL (served by serve.py),
never base64.note is what the user edits in the UI; PATCH writes the whole slides array back.<YTID>-slide-NN.jpg. Never reuse bare
slide-NN.jpg for a new video..md file is directly inside --dir (not a
subfolder) and the filename is <YTID>.md.VIDEO_LIBRARY_DIR) controls where the library lives.serve.py, stdlib + PyYAML) renders everything and handles
note write-back. Drop it anywhere Python runs.© dair-ai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 10 other files (scripts) in plugins/youtube-notetaker/skills/youtube-notetaker of dair-ai/dair-academy-plugins.
Open the folder on GitHubat commit 0abffdc
YouTube Talk Notetaker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| YouTube Talk Notetaker this skilldair-ai/dair-academy-plugins | 614 | — | ~2.3k | Automated safety check: Pass | MIT | |
| Video Lenskar2phi/video-lens | 113 | — | ~8.2k | Automated safety check: Notes | MIT | |
| Youtube Transcriptbrowser-act/skills | 6.1k | — | ~2.1k | Automated safety check: Pass | MIT | |
| URL to Markdown FetcherJimLiu/baoyu-skills | 27k | 1 repos | ~2.1k | Automated safety check: Pass | MIT | |
| Video To Notelike-attract/video-to-note | 124 | — | ~1k | Automated safety check: Pass | MIT | |
| Gbro Series Vocabpyang5166/gbro-series-vocab | 163 | — | ~1.1k | Automated safety check: Pass | MIT |
kar2phi/video-lens
Fetch a YouTube transcript and generate an executive summary, key points, and timestamped topic list as a polished HTML report.
browser-act/skills
YouTube transcript extraction and content reformatting: given a YouTube video URL, opens the video's transcript panel, extracts all timestamped segments, and transforms the raw transcript into…
JimLiu/baoyu-skills
Fetches a web page, X post, YouTube transcript or Hacker News thread through a Chrome-driven CLI and saves it as clean markdown.
like-attract/video-to-note
Generate structured, timestamped Markdown notes from videos (Bilibili, Douyin, YouTube, or local media files) using the local VideoToNo service.
pyang5166/gbro-series-vocab
追剧学英语 / Learn English vocabulary from TV series. An agent skill from pyang5166/gbro-series-vocab.
KIRVO-REPORTING/video-to-notes
Use immediately for any bare YouTube or YouTube Shorts URL, youtu.be link, Bilibili or b23.tv link, or other video URL; do not ask what the user wants.
dair-ai/dair-academy-plugins
Creates and maintains configurable research wikis: scaffold a folder, add sources, compile pages and indexes, and file query answers back.
dair-ai/dair-academy-plugins
Generates and edits images with Google's Gemini Nano Banana Pro model through the Gemini API, including photo edits and multi-image composition.
dair-ai/dair-academy-plugins
Builds a single-file HTML survey paper on an AI or ML topic from a research bundle the agent curates, with prose and SVG figures written by Kimi K2.6.
dair-ai/dair-academy-plugins
Help a user learn a topic through adaptive tutoring, lesson planning, practice, retrieval checks, explanations, study guides, or exercises.
dair-ai/dair-academy-plugins
Has several open-weight models answer a question, rank each other's anonymized answers, then lets a chairman model write the final response through Fireworks AI.
dair-ai/dair-academy-plugins
Builds a self-contained HTML digest of AI and agent news pulled from chosen X accounts through the official X MCP server, grouped into categories like Coding Agents and Agent Research.
Categories
Converts a YouTube talk into a markdown study note with slide images, a timestamped transcript and editable notes, browsable through a small local server. Each YouTube video you add becomes one plain markdown file in a local library folder, which defaults to ~/video-deepdives/ and can be changed with VIDEO_LIBRARY_DIR. The file holds video metadata and a slides list in its frontmatter and the full transcript as timestamped lines in the body, while slide images are saved in a media folder named per video.
YouTube Talk Notetaker fits situations like: studying a conference talk and capturing its slides with timestamps; building a personal library of YouTube talks as markdown files; taking searchable, timestamped notes on a video.
Run `npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a claude-code`. Or copy the skill folder (plugins/youtube-notetaker/skills/youtube-notetaker in dair-ai/dair-academy-plugins) into .claude/skills/youtube-notetaker in your project. Claude Code loads it when a task matches its description.
Run `npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a codex`. Or copy the skill folder (plugins/youtube-notetaker/skills/youtube-notetaker in dair-ai/dair-academy-plugins) into .agents/skills/youtube-notetaker in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add dair-ai/dair-academy-plugins --skill youtube-notetaker -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/youtube-notetaker, .gemini/skills/youtube-notetaker, .github/skills/youtube-notetaker and .opencode/skills/youtube-notetaker in your project.
Going by SKILL.md and its folder, YouTube Talk Notetaker needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3 and pip). Our summary lists: yt-dlp and ffmpeg on the PATH; Python 3 with Pillow and PyYAML.
SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
YouTube Talk Notetaker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with YouTube Talk Notetaker: Video Lens (kar2phi/video-lens, 113 stars), Youtube Transcript (browser-act/skills, 6.1k stars), URL to Markdown Fetcher (JimLiu/baoyu-skills, 27k stars) and Video To Note (like-attract/video-to-note, 124 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
dair-ai (a GitHub organization) maintains it in dair-ai/dair-academy-plugins, which has 614 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on July 21, 2026.
Source: dair-ai/dair-academy-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.