Agent skill

Multimodal Extraction

by swyxio in swyxio/skills

Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…

MITAuto-check passedDocuments & Office

Install Multimodal Extraction

skills CLI
$ npx skills add swyxio/skills --skill multimodal-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills multimodal-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/multimodal-extraction .claude/skills/multimodal-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
multimodal-extraction
GitHub stars
175
Token cost
~922 tokens
SKILL.md length
316 words
Files
3
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…

  • Works in 4 steps: Resolve the Source → Extract Visual Anchors → Transcribe → …
  • Asked to turn a video into a multimodal notes file
  • SKILL.md covers Overview, When To Use, Requirements and Command, plus 4 more sections
  • Runs Python scripts from its folder; calls python3, brew and pip3

What it does

Multimodal Extraction is an agent skill from swyxio/skills. Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the transcript at the associated timestamps. Use when asked to turn a video into a multimodal notes file, slide-synced transcript, screenshot-enhanced transcript, or talk recap with images.

Its SKILL.md is about 920 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `README.md` and `multimodal_extract.py`).

It sits in Documents & Office, covering Transcription, Slides and decks and Markdown. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • Asked to turn a video into a multimodal notes file
  • Slide-synced transcript
  • Screenshot-enhanced transcript
  • Talk recap with images

Example prompts

  • “/multimodal-extraction”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Resolve the Source
  2. Extract Visual Anchors
  3. Transcribe
  4. Merge into Markdown

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • brew
    • pip3
    • ffmpeg
    • whisper

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip3, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Multimodal Extraction loads about 922 tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 316 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~922

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 316 words, ~922 tokens.

Download SKILL.mdSave it as .claude/skills/multimodal-extraction/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
multimodal-extraction
description
Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the transcript at the associated timestamps. Use when asked to turn a video into a multimodal notes file, slide-synced transcript, screenshot-enhanced transcript, or talk recap with images.
metadata.version
0.1.0

Multimodal Extraction

Overview

This skill composes the existing video workflows into one artifact:

  • download-video for URL inputs
  • thumbnail-extraction for slide frames and key screenshots
  • transcribe-anything for transcript strategy

The implementation is intentionally speed-first:

  1. Download only when the input is a URL
  2. Reuse the fast slide/key-frame heuristics from thumbnail-extraction
  3. Use local whisper JSON output for timestamped transcript segments
  4. Merge everything into one Markdown timeline with relative image links

When To Use

  • "Turn this talk into multimodal notes"
  • "Make me a markdown transcript with screenshots"
  • "Extract slides and transcript together"
  • "Build a recap doc from this video"
  • "Given this YouTube URL, produce a slide-synced transcript"

Requirements

bash
brew install ffmpeg yt-dlp
pip3 install --break-system-packages openai-whisper

The following existing local script is reused:

  • ../thumbnail-extraction/thumbnail_extractor.py

Command

bash
python3 multimodal_extract.py <video_or_url> [output_dir] [--language en] [--whisper-model turbo] [--top-n 4]

What It Does

Step 1: Resolve the Source
  • If the input is a local file, use it directly
  • If the input starts with http:// or https://, download it first with yt-dlp
  • For YouTube URLs, direct yt-dlp is usually enough
  • For trickier hosted pages, this skill follows the same practical intent as download-video: get a usable local file first
Step 2: Extract Visual Anchors

Run:

bash
python3 ../thumbnail-extraction/thumbnail_extractor.py "$VIDEO" "$OUTPUT/visuals" 4 --extract-slides

This produces:

  • top thumbnail candidates in the root of visuals/
  • slide images in visuals/slides/
  • manifests with timestamps
Step 3: Transcribe

Extract normalized mono 16k audio:

bash
ffmpeg -y -i "$VIDEO" -vn -ac 1 -ar 16000 -acodec pcm_s16le \
  -af "highpass=f=80,lowpass=f=8000,loudnorm=I=-16:TP=-1.5:LRA=11" \
  "$OUTPUT/audio/source_preprocessed.wav"

Then transcribe with Whisper:

bash
whisper "$OUTPUT/audio/source_preprocessed.wav" \
  --model turbo \
  --language en \
  --word_timestamps True \
  --condition_on_previous_text False \
  --output_format json \
  --output_dir "$OUTPUT/transcript"
Step 4: Merge into Markdown

The script:

  • reads slide and thumbnail manifests
  • reads Whisper transcript segments
  • sorts all visual anchors by timestamp
  • groups transcript text between successive visual anchors
  • writes multimodal_timeline.md with:
    • section timestamp
    • associated image(s)
    • transcript span for that interval

Output

output_dir/
  source/
  visuals/
  audio/
  transcript/
  multimodal_timeline.md

Design Principle

The goal is total end-to-end extraction speed.

That means:

  • heuristics first
  • local transcript by default
  • no VLM in the common path
  • only enough structure to make the Markdown artifact useful immediately

Future Extensions

  • add backend switching for transcribe-anything
  • add deck-aware slide labeling when a source deck exists
  • add speaker diarization sections
  • add chaptering or summary generation on top of the Markdown timeline

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in multimodal-extraction of swyxio/skills.

  • SKILL.md
  • README.md
  • multimodal_extract.py

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Multimodal Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Multimodal Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Multimodal Extraction this skillswyxio/skills175—~922Automated safety check: PassMIT
Youtube Notetakersickn33/agentic-awesome-skills47k1 repos~2.3kAutomated safety check: PassMIT
Human Reviewpetergyang/human-review1.4k—~2.1kAutomated safety check: PassMIT
Save Mdmblode/agent-skills144—~537Automated safety check: PassMIT
MarkitdownImCa0/just-laws78214 repos~3.2kAutomated safety check: NotesMIT
Suikonwiizo/suiko114—~1.8kAutomated safety check: PassMIT

Similar skills

  • Youtube Notetaker

    sickn33/agentic-awesome-skills

    Turn YouTube talks into local study notes with slides, transcripts, editable annotations, and a markdown-backed viewer.

    47k GitHub starsUsed in 1 repo~2.3k tokens
    Documents & OfficeAuto-check passed
  • Human Review

    petergyang/human-review

    Open an HTML file, Markdown file, or localhost page in the browser so the user can edit text directly and leave comments on specific parts, then send all edits and comments back to you.

    1.4k GitHub stars~2.1k tokensUpdated 22 days ago
    Documents & OfficeAuto-check passed
  • Save Md

    mblode/agent-skills

    Saves a named source to Markdown with provenance and faithful extraction through direct export endpoints.

    144 GitHub stars~537 tokensUpdated yesterday
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    782 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Suiko

    nwiizo/suiko

    日本語文書のAI由来の均一さ、翻訳調、不自然さ、論旨、読解負荷を、決定的なRust CLIと目視で診断し、依頼に応じて書く・直す。日本語の学術論文・研究報告では、中心命題、用語、論証、DOCX/PDF納品を監査契約で確認する。Use when the user explicitly mentions suiko, asks whether Japanese text looks…

    114 GitHub stars~1.8k tokensUpdated today
    Documents & OfficeAuto-check passed
  • Ky Markdown Rebuilder

    KyrieCheungYep/ky-markdown-rebuilder

    Rebuild visual documents into reliable Markdown by combining text extraction with page or screenshot alignment.

    117 GitHub stars~5.7k tokensUpdated 3 mo ago
    Documents & OfficeAuto-check passed

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    175 GitHub stars~1.1k tokensUpdated 4 days ago
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    175 GitHub stars~4.3k tokensUpdated 4 days ago
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    175 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    175 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed

Questions about Multimodal Extraction

What does Multimodal Extraction do?

Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the…. Multimodal Extraction is an agent skill from swyxio/skills. Given a local video or video URL, downloads the media if needed, extracts slide frames and key moments, transcribes the audio, and writes a Markdown timeline that interleaves screenshots with the transcript at the associated timestamps.

When should I use Multimodal Extraction?

Multimodal Extraction fits situations like: asked to turn a video into a multimodal notes file; slide-synced transcript; screenshot-enhanced transcript; talk recap with images.

How do I install Multimodal Extraction in Claude Code?

Run `npx skills add swyxio/skills --skill multimodal-extraction -a claude-code`. Or copy the skill folder (multimodal-extraction in swyxio/skills) into .claude/skills/multimodal-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Multimodal Extraction in Codex?

Run `npx skills add swyxio/skills --skill multimodal-extraction -a codex`. Or copy the skill folder (multimodal-extraction in swyxio/skills) into .agents/skills/multimodal-extraction in your project. Codex loads it when a task matches its description.

Can I use Multimodal Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill multimodal-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/multimodal-extraction, .gemini/skills/multimodal-extraction, .github/skills/multimodal-extraction and .opencode/skills/multimodal-extraction in your project.

What does Multimodal Extraction need to run?

Going by SKILL.md and its folder, Multimodal Extraction needs Python for the scripts in its folder and the command-line tools its instructions call (python3, brew, pip3, ffmpeg and whisper). Our summary lists: Python 3.

Does Multimodal Extraction access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Multimodal Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Multimodal Extraction use?

Multimodal Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Multimodal Extraction use?

About 922 tokens (SKILL.md is roughly 3.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Multimodal Extraction?

Skills that share tags, products or a category with Multimodal Extraction: Youtube Notetaker (sickn33/agentic-awesome-skills, 47k stars), Human Review (petergyang/human-review, 1.4k stars), Save Md (mblode/agent-skills, 144 stars) and Markitdown (ImCa0/just-laws, 782 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Multimodal Extraction?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 175 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.