Agent skill

Capture Video Frames

by pamelafox in pamelafox/presentation-skills

Capture frames from a YouTube video at a regular interval, produce a manifest mapping filenames to timestamps, and describe each frame with an LLM.

MITAuto-check passedAgent Workflows

Install Capture Video Frames

skills CLI
$ npx skills add pamelafox/presentation-skills --skill capture-video-frames -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pamelafox/presentation-skills capture-video-frames --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pamelafox/presentation-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/capture-video-frames .claude/skills/capture-video-frames && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
capture-video-frames
GitHub stars
125
Token cost
~1.7k tokens
SKILL.md length
663 words
Files
2
Skills in repo
14
Repo updated
First seen
Licence
MIT

At a glance

Capture frames from a YouTube video at a regular interval, produce a manifest mapping filenames to timestamps, and describe each frame with an LLM.

  • Works in 3 steps: Capture frames → Describe frames using the describe-frame… → Deduplicate frames and select best…
  • : extract video frames
  • SKILL.md covers Step 1: Capture frames, Step 2: Describe frames using… and Step 3: Deduplicate frames and…
  • Runs Python scripts from its folder; calls brew, uv and yt-dlp

What it does

Capture Video Frames is an agent skill from pamelafox/presentation-skills. Capture frames from a YouTube video at a regular interval, produce a manifest mapping filenames to timestamps, and describe each frame with an LLM. USE FOR: extract video frames, capture screenshots from YouTube, describe video frames, video frame analysis, frame-by-frame summary.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `capture_video_frames.py`).

It sits in Agent Workflows. It works with YouTube and FFmpeg. The repository describes itself as: Skills for AI agents to process presentations - helpful for teachers and speakers. The licence is MIT.

When your agent uses it

  • : extract video frames
  • Capture screenshots from YouTube
  • Describe video frames
  • Video frame analysis

Example prompts

  • “/capture-video-frames”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Capture frames
  2. Describe frames using the describe-frame subagent
  3. Deduplicate frames and select best speaker faces

What it can do on your machine

Read from SKILL.md and the folder at commit 2b809b3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • brew
    • uv
    • yt-dlp
    • ffmpeg
    • pip
    • apt-get

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Capture Video Frames loads about 1.7k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 663 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from pamelafox/presentation-skills at commit 2b809b3, republished under its MIT licence (© pamelafox). 663 words, ~1,689 tokens.

Download SKILL.mdSave it as .claude/skills/capture-video-frames/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
capture-video-frames
description
Capture frames from a YouTube video at a regular interval, produce a manifest mapping filenames to timestamps, and describe each frame with an LLM. USE FOR: extract video frames, capture screenshots from YouTube, describe video frames, video frame analysis, frame-by-frame summary.
argument-hint
<youtube_url> <output_dir> [--interval SECONDS]

Capture and describe video frames

Step 1: Capture frames

Run the capture_video_frames.py script:

bash
uv run .agents/skills/capture-video-frames/capture_video_frames.py <youtube_url> <output_dir> [--interval SECONDS]
Arguments
  • youtube_url (required): YouTube video URL (same formats accepted by the extract-transcript skill).
  • output_dir (required): Directory to save frames and the manifest file. Created if it doesn't exist.
  • --interval (optional): Seconds between captured frames. Defaults to 30.
Outputs
  • frame_0000.png, frame_0030.png, … — PNG images named by their timestamp in seconds (zero-padded to 4 digits).
  • frames_manifest.md — A markdown file listing each frame with its timestamp and a placeholder for descriptions.

Example frames_manifest.md:

| File | Timestamp | Description |
|------|-----------|-------------|
| frame_0000.png | [00:00] | |
| frame_0030.png | [00:30] | |
| frame_0060.png | [01:00] | |
Prerequisites
  • yt-dlp: brew install yt-dlp or pip install yt-dlp
  • ffmpeg: brew install ffmpeg or apt-get install ffmpeg

Step 2: Describe frames using the describe-frame subagent

After capturing frames, describe each frame by running the describe-frame custom agent as a subagent. Each subagent invocation gets an isolated context, so frame images won't accumulate and exhaust the context window.

The describe-frame agent is defined in .github/agents/describe-frame.md.

Procedure
  1. Read frames_manifest.md from the output directory to get the full list of frames.
  2. For each frame, run the describe-frame agent as a subagent with a prompt that includes:
    • The absolute path to the current frame image to view.
    • The absolute path to the previous frame image to view (if one exists).
    • The previous frame's description as text (if one exists).
  3. The subagent will return a plain-text description (or (same as previous) if the frame is essentially identical to the previous one).
  4. After each subagent returns, update the Description column for that row in frames_manifest.md immediately.
  5. Continue until all frames are described.
Subagent prompt template

Use this as the prompt when invoking the describe-frame subagent (fill in the bracketed values):

Describe the current frame image at: [CURRENT_FRAME_ABSOLUTE_PATH]

[If previous frame exists, include these two lines:]
The previous frame image is at: [PREVIOUS_FRAME_ABSOLUTE_PATH]
The previous frame was described as: "[PREVIOUS_DESCRIPTION]"
Example output

After describing all frames, frames_manifest.md should look like:

| File | Timestamp | Description |
|------|-----------|-------------|
| frame_0000.png | [00:00] | Title slide introducing "Building RAG apps with Python" |
| frame_0030.png | [00:30] | Speaker showing the agenda with four main topics |
| frame_0060.png | [01:00] | (same as previous) |
| frame_0090.png | [01:30] | Architecture diagram of a retrieval-augmented generation pipeline |

Step 3: Deduplicate frames and select best speaker faces

After all frames are described, groups of consecutive (same as previous) rows represent the same visual content captured at different moments. Within each group, speaker faces may differ — eyes open vs closed, mouth open vs closed, facing camera vs turned away.

Goal

For each group of duplicate frames, keep only one frame — the one with the best speaker face quality — and remove the rest.

Criteria for best face (in priority order)
  1. The current speaker's mouth should be open (mid-speech). If you know who is speaking at that timestamp (from a transcript), prioritize that speaker.
  2. Eyes open — no mid-blink frames.
  3. Facing camera — not turned sideways or looking down.
  4. If no speakers are visible (e.g., full-screen demo or slide without webcam feeds), all frames in the group are equivalent — keep the first one.
Show full SKILL.md (258 more words)Show less
Procedure
  1. Identify all groups of consecutive rows where the description is (same as previous). Each group starts with the "anchor" frame (the one with an actual description) followed by one or more (same as previous) rows.
  2. For each group, run the describe-frame subagent to compare faces across the anchor frame and each duplicate. Use this prompt template:
Compare these two frames focusing ONLY on the speaker faces visible in webcam feeds. Which frame has better speaker faces — eyes open, facing camera, mouth open (mid-speech), not mid-blink or turned away?

Frame A: [ANCHOR_FRAME_ABSOLUTE_PATH]
Frame B: [DUPLICATE_FRAME_ABSOLUTE_PATH]

Reply with ONLY one of:
- "A BETTER" if the anchor frame has better speaker faces
- "B BETTER" if the duplicate frame has better speaker faces
- "EQUAL" if both are equivalent
- "NO SPEAKERS" if no speaker faces are visible in either frame

Then add a brief reason.
  1. After comparing all duplicates in a group against the anchor (and the current best), determine the single best frame.
  2. If the best frame is NOT the anchor:
    • Move the anchor's description to the best frame's row.
    • Add a face-quality note to the description, e.g., (better speaker faces than frame_XXXX: eyes open, facing camera)
  3. Remove all other (same as previous) rows from the manifest.
  4. If the anchor was already the best, just remove the duplicate rows.
Recapturing frames for closed mouths

After deduplication, if the best frame in a group still has the speaking person's mouth closed (both speakers have mouths closed), try recapturing at nearby timestamps:

  1. Download the video if not already available:
    bash
    yt-dlp -f "bestvideo[height<=720]" --no-playlist -o "<output_dir>/video.%(ext)s" "<youtube_url>"
  2. Capture alternative frames at +2s, +5s, +8s, and +10s offsets from the frame's timestamp:
    bash
    ffmpeg -ss <SECONDS> -i <output_dir>/video.mp4 -frames:v 1 -q:v 2 <output_dir>/alt_<FRAME>_<SECONDS>.png -y
  3. Use the describe-frame subagent to check if the speaker's mouth is open in any alternative, AND that the slide/demo content is still the same.
  4. If a better alternative is found, replace the frame file (cp alt_XXXX.png frame_XXXX.png).
  5. Clean up: rm -f <output_dir>/alt_*.png <output_dir>/video.mp4

© pamelafox, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/capture-video-frames of pamelafox/presentation-skills.

  • SKILL.md
  • capture_video_frames.py

Open the folder on GitHubat commit 2b809b3

Compare with similar skills

Capture Video Frames next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Capture Video Frames compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Capture Video Frames this skillpamelafox/presentation-skills125—~1.7kAutomated safety check: PassMIT
Youtube Clipperop7418/Youtube-clipper-skill2.2k—~1.6kAutomated safety check: NotesMIT
Video Transcribewendy7756/AI-Video-Transcriber3.3k—~937Automated safety check: NotesApache-2.0
Claude Real VideoHUANGCHIHHUNGLeo/claude-real-video2.2k—~639Automated safety check: PassMIT
Ffmpeg Skillkajisho5/ffmpeg-skill1.9k—~7.4kAutomated safety check: PassMIT
Video Perceptionjordanrendric/claude-video-vision1.4k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Youtube Clipper

    op7418/Youtube-clipper-skill

    YouTube 视频智能剪辑工具。下载视频和字幕,AI 分析生成精细章节(几分钟级别), 用户选择片段后自动剪辑、翻译字幕为中英双语、烧录字幕到视频,并生成总结文案。

    2.2k GitHub stars~1.6k tokensUpdated 8 mo ago
    Media & CreativeAuto-check: notes
  • Video Transcribe

    wendy7756/AI-Video-Transcriber

    Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.

    3.3k GitHub stars~937 tokensUpdated 23 days ago
    Media & CreativeAuto-check: notes
  • Claude Real Video

    HUANGCHIHHUNGLeo/claude-real-video

    Watch a video for the user. An agent skill from HUANGCHIHHUNGLeo/claude-real-video.

    2.2k GitHub stars~639 tokensUpdated 6 days ago
    Media & CreativeAuto-check passed
  • Ffmpeg Skill

    kajisho5/ffmpeg-skill

    Edit video and audio with local FFmpeg from natural-language requests: cut, trim, join, resize/reframe (9:16, 1:1), speed change, captions and subtitles (SRT/ASS, animated, karaoke), logos and text…

    1.9k GitHub stars~7.4k tokensUpdated 3 days ago
    Media & CreativeAuto-check passed
  • Video Perception

    jordanrendric/claude-video-vision

    A skill your agent uses when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation

    1.4k GitHub stars~1.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Explainroo

    vincentsch/explainroo

    Make an explainer video (MP4) with a voice-over using explainroo.

    489 GitHub stars~439 tokensUpdated 3 days ago
    Media & CreativeAuto-check passed

More from pamelafox/presentation-skills

All 14 skills in this repo
  • Generate Images Mai

    pamelafox/presentation-skills

    Generate or edit bitmap images with Microsoft MAI-Image-2.5 through the Azure AI image APIs.

    125 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check: notes
  • Make Revealjs Presentation

    pamelafox/presentation-skills

    Create or update a RevealJS HTML presentation using the repository's bundled slide template.

    125 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Fetch Slides

    pamelafox/presentation-skills

    Fetch presentation slides from a URL and convert them to PDF.

    125 GitHub stars~528 tokensUpdated 1 mo ago
    Auto-check passed
  • Generate Writeup

    pamelafox/presentation-skills

    Generate an annotated blog-style write-up from a presentation's slides and video recording.

    125 GitHub stars~2.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Outline Slides

    pamelafox/presentation-skills

    Generate a numbered outline of presentation slides with one-sentence summaries.

    125 GitHub stars~462 tokensUpdated 1 mo ago
    Auto-check passed
  • Thumbnail Of PPTX

    pamelafox/presentation-skills

    Capture a thumbnail image of a slide from a OneDrive/Office presentation link.

    125 GitHub stars~686 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Categories

Questions about Capture Video Frames

What does Capture Video Frames do?

Capture frames from a YouTube video at a regular interval, produce a manifest mapping filenames to timestamps, and describe each frame with an LLM. Capture Video Frames is an agent skill from pamelafox/presentation-skills. Capture frames from a YouTube video at a regular interval, produce a manifest mapping filenames to timestamps, and describe each frame with an LLM.

When should I use Capture Video Frames?

Capture Video Frames fits situations like: : extract video frames; capture screenshots from YouTube; describe video frames; video frame analysis.

How do I install Capture Video Frames in Claude Code?

Run `npx skills add pamelafox/presentation-skills --skill capture-video-frames -a claude-code`. Or copy the skill folder (.agents/skills/capture-video-frames in pamelafox/presentation-skills) into .claude/skills/capture-video-frames in your project. Claude Code loads it when a task matches its description.

How do I install Capture Video Frames in Codex?

Run `npx skills add pamelafox/presentation-skills --skill capture-video-frames -a codex`. Or copy the skill folder (.agents/skills/capture-video-frames in pamelafox/presentation-skills) into .agents/skills/capture-video-frames in your project. Codex loads it when a task matches its description.

Can I use Capture Video Frames in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pamelafox/presentation-skills --skill capture-video-frames -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/capture-video-frames, .gemini/skills/capture-video-frames, .github/skills/capture-video-frames and .opencode/skills/capture-video-frames in your project.

What does Capture Video Frames need to run?

Going by SKILL.md and its folder, Capture Video Frames needs Python for the scripts in its folder and the command-line tools its instructions call (brew, uv, yt-dlp, ffmpeg, pip and apt-get). Our summary lists: Python 3.

Does Capture Video Frames access the network?

SKILL.md contains no URLs. Its commands use uv and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Capture Video Frames safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Capture Video Frames use?

Capture Video Frames is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Capture Video Frames use?

About 1.7k tokens (SKILL.md is roughly 6.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Capture Video Frames?

Skills that share tags, products or a category with Capture Video Frames: Youtube Clipper (op7418/Youtube-clipper-skill, 2.2k stars), Video Transcribe (wendy7756/AI-Video-Transcriber, 3.3k stars), Claude Real Video (HUANGCHIHHUNGLeo/claude-real-video, 2.2k stars) and Ffmpeg Skill (kajisho5/ffmpeg-skill, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Capture Video Frames?

pamelafox (a GitHub user) maintains it in pamelafox/presentation-skills, which has 125 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on September 2, 2026.

Source: pamelafox/presentation-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.