Youtube Notetaker
sickn33/agentic-awesome-skills
Turn YouTube talks into local study notes with slides, transcripts, editable annotations, and a markdown-backed viewer.
Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…
$ npx skills add swyxio/skills --skill multimodal-extraction -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install swyxio/skills multimodal-extraction --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/multimodal-extraction .claude/skills/multimodal-extraction && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "multimodal-extraction" agent skill from https://github.com/swyxio/skills/tree/main/multimodal-extraction into .claude/skills/multimodal-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multimodal-extraction", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/swyxio/skills/tree/main/multimodal-extractionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add swyxio/skills --skill multimodal-extraction -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install swyxio/skills multimodal-extraction --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/multimodal-extraction .agents/skills/multimodal-extraction && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "multimodal-extraction" agent skill from https://github.com/swyxio/skills/tree/main/multimodal-extraction into .agents/skills/multimodal-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multimodal-extraction", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill multimodal-extraction -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install swyxio/skills multimodal-extraction --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/multimodal-extraction .cursor/skills/multimodal-extraction && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "multimodal-extraction" agent skill from https://github.com/swyxio/skills/tree/main/multimodal-extraction into .cursor/skills/multimodal-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multimodal-extraction", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/swyxio/skills.git --path multimodal-extraction--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add swyxio/skills --skill multimodal-extraction -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install swyxio/skills multimodal-extraction --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/multimodal-extraction .gemini/skills/multimodal-extraction && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "multimodal-extraction" agent skill from https://github.com/swyxio/skills/tree/main/multimodal-extraction into .gemini/skills/multimodal-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multimodal-extraction", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install swyxio/skills multimodal-extractionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add swyxio/skills --skill multimodal-extraction -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/multimodal-extraction .github/skills/multimodal-extraction && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "multimodal-extraction" agent skill from https://github.com/swyxio/skills/tree/main/multimodal-extraction into .github/skills/multimodal-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multimodal-extraction", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill multimodal-extraction -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install swyxio/skills multimodal-extraction --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/multimodal-extraction .opencode/skills/multimodal-extraction && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "multimodal-extraction" agent skill from https://github.com/swyxio/skills/tree/main/multimodal-extraction into .opencode/skills/multimodal-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "multimodal-extraction", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
multimodal-extractionGiven a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…
Multimodal Extraction is an agent skill from swyxio/skills. Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the transcript at the associated timestamps. Use when asked to turn a video into a multimodal notes file, slide-synced transcript, screenshot-enhanced transcript, or talk recap with images.
Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `README.md` and `multimodal_extract.py`).
It sits in Documents & Office, covering Transcription, Slides and decks and Markdown. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
python3brewpip3ffmpegwhisperFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip3, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Multimodal Extraction loads about 922 tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 316 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 316 words, ~922 tokens.
.claude/skills/multimodal-extraction/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.This skill composes the existing video workflows into one artifact:
download-video for URL inputsthumbnail-extraction for slide frames and key screenshotstranscribe-anything for transcript strategyThe implementation is intentionally speed-first:
thumbnail-extractionwhisper JSON output for timestamped transcript segmentsbrew install ffmpeg yt-dlp
pip3 install --break-system-packages openai-whisperThe following existing local script is reused:
../thumbnail-extraction/thumbnail_extractor.pypython3 multimodal_extract.py <video_or_url> [output_dir] [--language en] [--whisper-model turbo] [--top-n 4]http:// or https://, download it first with yt-dlpyt-dlp is usually enoughdownload-video: get a usable local file firstRun:
python3 ../thumbnail-extraction/thumbnail_extractor.py "$VIDEO" "$OUTPUT/visuals" 4 --extract-slidesThis produces:
visuals/visuals/slides/Extract normalized mono 16k audio:
ffmpeg -y -i "$VIDEO" -vn -ac 1 -ar 16000 -acodec pcm_s16le \
-af "highpass=f=80,lowpass=f=8000,loudnorm=I=-16:TP=-1.5:LRA=11" \
"$OUTPUT/audio/source_preprocessed.wav"Then transcribe with Whisper:
whisper "$OUTPUT/audio/source_preprocessed.wav" \
--model turbo \
--language en \
--word_timestamps True \
--condition_on_previous_text False \
--output_format json \
--output_dir "$OUTPUT/transcript"The script:
multimodal_timeline.md with:output_dir/
source/
visuals/
audio/
transcript/
multimodal_timeline.mdThe goal is total end-to-end extraction speed.
That means:
transcribe-anything© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in multimodal-extraction of swyxio/skills.
Open the folder on GitHubat commit 038ef34
Multimodal Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Multimodal Extraction this skillswyxio/skills | 175 | — | ~922 | Automated safety check: Pass | MIT | |
| Youtube Notetakersickn33/agentic-awesome-skills | 47k | 1 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Human Reviewpetergyang/human-review | 1.4k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Save Mdmblode/agent-skills | 144 | — | ~537 | Automated safety check: Pass | MIT | |
| MarkitdownImCa0/just-laws | 782 | 14 repos | ~3.2k | Automated safety check: Notes | MIT | |
| Suikonwiizo/suiko | 114 | — | ~1.8k | Automated safety check: Pass | MIT |
sickn33/agentic-awesome-skills
Turn YouTube talks into local study notes with slides, transcripts, editable annotations, and a markdown-backed viewer.
petergyang/human-review
Open an HTML file, Markdown file, or localhost page in the browser so the user can edit text directly and leave comments on specific parts, then send all edits and comments back to you.
mblode/agent-skills
Saves a named source to Markdown with provenance and faithful extraction through direct export endpoints.
ImCa0/just-laws
Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.
nwiizo/suiko
日本語文書のAI由来の均一さ、翻訳調、不自然さ、論旨、読解負荷を、決定的なRust CLIと目視で診断し、依頼に応じて書く・直す。日本語の学術論文・研究報告では、中心命題、用語、論証、DOCX/PDF納品を監査契約で確認する。Use when the user explicitly mentions suiko, asks whether Japanese text looks…
KyrieCheungYep/ky-markdown-rebuilder
Rebuild visual documents into reliable Markdown by combining text extraction with page or screenshot alignment.
swyxio/skills
Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.
swyxio/skills
Design, implement, audit, or refresh protected username and handle namespaces for public products.
swyxio/skills
Fully automated new Mac setup for fullstack web developers and AI engineers.
swyxio/skills
Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…
swyxio/skills
Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.
swyxio/skills
Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.
Categories
Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…. Multimodal Extraction is an agent skill from swyxio/skills. Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the transcript at the associated timestamps.
Multimodal Extraction fits situations like: asked to turn a video into a multimodal notes file; slide-synced transcript; screenshot-enhanced transcript; talk recap with images.
Run `npx skills add swyxio/skills --skill multimodal-extraction -a claude-code`. Or copy the skill folder (multimodal-extraction in swyxio/skills) into .claude/skills/multimodal-extraction in your project. Claude Code loads it when a task matches its description.
Run `npx skills add swyxio/skills --skill multimodal-extraction -a codex`. Or copy the skill folder (multimodal-extraction in swyxio/skills) into .agents/skills/multimodal-extraction in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill multimodal-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multimodal-extraction, .gemini/skills/multimodal-extraction, .github/skills/multimodal-extraction and .opencode/skills/multimodal-extraction in your project.
Going by SKILL.md and its folder, Multimodal Extraction needs Python for the scripts in its folder and the command-line tools its instructions call (python3, brew, pip3, ffmpeg and whisper). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Multimodal Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 922 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Multimodal Extraction: Youtube Notetaker (sickn33/agentic-awesome-skills, 47k stars), Human Review (petergyang/human-review, 1.4k stars), Save Md (mblode/agent-skills, 144 stars) and Markitdown (ImCa0/just-laws, 782 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
swyxio (a GitHub user) maintains it in swyxio/skills, which has 175 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.
Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.