Agent skill

Thumbnail Extraction

by swyxio in swyxio/skills

Extracts the most interesting frames from video files for thumbnail compositing.

MITAuto-check passedMedia & Creative

Install Thumbnail Extraction

skills CLI
$ npx skills add swyxio/skills --skill thumbnail-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install swyxio/skills thumbnail-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/thumbnail-extraction .claude/skills/thumbnail-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
thumbnail-extraction
GitHub stars
175
Token cost
~2.4k tokens
SKILL.md length
917 words
Files
3
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Extracts the most interesting frames from video files for thumbnail compositing.

  • Works in 3 steps: Quadrant selection (Pass 1 → Pass 2):… → Segment-forced selection (Pass 2 →… → Fallback: If any segment is empty, fill…
  • Asked to extract thumbnails
  • SKILL.md covers Overview, When to Use, Dependencies and Pipeline Architecture, plus 5 more sections
  • Runs Python scripts from its folder; calls python3, pip and pip3; reaches youtube.com

What it does

Thumbnail Extraction is an agent skill from swyxio/skills. Extracts the most interesting frames from video files for thumbnail compositing. Detects faces, expressions, smiles, and presentation slides. Outputs full frames, face crops, and transparent cutouts. Use when asked to extract thumbnails, find interesting frames, grab screenshots from video, or create thumbnail candidates from recordings.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `README.md` and `thumbnail_extractor.py`).

It sits in Media & Creative, covering Drug discovery and cheminformatics and Slides and decks. It works with YouTube. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.

When your agent uses it

  • Asked to extract thumbnails
  • Find interesting frames
  • Grab screenshots from video
  • Create thumbnail candidates from recordings

Example prompts

  • “Use the thumbnail-extraction skill to extract the most interesting frames from video files for thumbnail compositing”
  • “/thumbnail-extraction”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Quadrant selection (Pass 1 → Pass 2): Divide video duration into N segments, pick the highest-scoring frame from each segment
  2. Segment-forced selection (Pass 2 → Final): Divide top candidates into top_n equal time segments, pick best from each
  3. Fallback: If any segment is empty, fill from overall top scores

What it can do on your machine

Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • pip
    • pip3
    • yt-dlp

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Thumbnail Extraction loads about 2.4k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 917 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~90
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 917 words, ~2,392 tokens.

Download SKILL.mdSave it as .claude/skills/thumbnail-extraction/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
thumbnail-extraction
description
Extracts the most interesting frames from video files for thumbnail compositing. Detects faces, expressions, smiles, and presentation slides. Outputs full frames, face crops, and transparent cutouts. Use when asked to extract thumbnails, find interesting frames, grab screenshots from video, or create thumbnail candidates from recordings.
metadata.version
0.1.0

Video Thumbnail Extraction

Overview

Automatically scan a local MP4 video (or YouTube URL via yt-dlp) and extract the 4 most visually interesting frames — prioritizing expressive faces (laughing, shocked, smiling) and engaging presentation slides. Outputs full frames, face crops, and background-removed transparent PNGs ready for compositing.

When to Use

  • Before creating YouTube thumbnails (feeds into youtube-thumbnails skill)
  • When you need the best screenshot from a long video recording
  • When compositing a thumbnail and need transparent guest cutouts
  • Processing Zoom gallery recordings, interviews, or presentations

Dependencies

Python packages (install once)
bash
# In sandbox (Cowork VM):
pip install opencv-python scenedetect deepface pillow numpy --break-system-packages

# On host Mac (for background removal — sandbox can't download the model):
pip3 install 'rembg[cpu]' pillow --break-system-packages
System tools
  • ffmpeg (usually pre-installed)
  • python3 (3.10+)
  • yt-dlp (optional, for YouTube URLs): pip install yt-dlp --break-system-packages
Model downloads (first run)
  • DeepFace expression model (~1MB): Downloads automatically on first use. If blocked by proxy, expression detection falls back to OpenCV smile cascade (still effective).
  • rembg u2net model (~176MB): Downloads on first use. Must run on host Mac if sandbox blocks GitHub releases.

Pipeline Architecture

Two-Pass Design (memory-efficient)

Pass 1 — Quick Scan (OpenCV only, no deep learning)

  • Sample a frame every 10 seconds across the video
  • Skip first/last 60 seconds (intro/outro)
  • For each frame:
    • Detect faces via Haar cascade (fast, no GPU needed)
    • Detect smiles within face regions
    • Compute visual variance (proxy for "interesting" content)
    • Detect presentation slides (high edge density + low color saturation)
  • Score each frame based on: face count, smile count, smile size, visual variance
  • Select top 12 diverse candidates using quadrant system: divide video into N time segments, pick best from each → ensures temporal spread

Pass 2 — Deep Analysis (DeepFace, only on top 12 candidates)

  • Re-read only the selected frames from video
  • Run DeepFace emotion detection (happy, surprise, fear, sad, angry, disgust, neutral)
  • Weight emotions by thumbnail value: happy > surprise > fear > angry > sad > neutral
  • Combine Pass 1 score with expression score
  • Final selection: divide candidates into N time segments, pick best from each → guarantees spread across the full video

Pass 3 — Output (rembg, only on final 4 frames)

  • Save full frame as JPG (95% quality)
  • Crop largest detected face with generous padding (0.5x)
  • Run background removal on face crop → transparent PNG
  • Generate manifest JSON with metadata
Scoring Heuristics
SignalWeightNotes
Face detected+2.0 per face (cap 3)Gallery views score high
Smile detected+3.0 per smileCascade-based, no model needed
Smile size ratio+5.0 × ratioBigger smiles = more expressive
Multi-person shot+1.0 bonus2+ faces = engaging
Happy expression+2.0 bonus (Pass 2)Best for thumbnails
Surprise expression+2.0 bonus (Pass 2)Eye-catching
Fear/angry expression+1.0 bonus (Pass 2)"Shocked" reactions
Visual variance+0.0–1.5Normalized by frame complexity
Presentation slidebaseline 1.5Useful for slide screenshots
Temporal Diversity Algorithm

The pipeline enforces temporal spread to avoid clustering picks in one segment:

  1. Quadrant selection (Pass 1 → Pass 2): Divide video duration into N segments, pick the highest-scoring frame from each segment
  2. Segment-forced selection (Pass 2 → Final): Divide top candidates into top_n equal time segments, pick best from each
  3. Fallback: If any segment is empty, fill from overall top scores

This ensures a 76-minute video yields picks from different parts (e.g., 1:00, 2:10, 21:50, 48:50) rather than clustering in the most face-heavy section.

Usage

Command Line
bash
python3 thumbnail_extractor.py <video_path> [output_dir] [top_n]

Arguments:

  • video_path — Path to MP4 file (required)
  • output_dir — Where to save outputs (default: ~/Downloads/thumb_candidates)
  • top_n — Number of candidates to extract (default: 4)

Examples:

bash
# Basic — extract 4 best frames
python3 thumbnail_extractor.py "GMT20260130-210038_Recording_gallery_2380x1544.mp4"

# Custom output dir and count
python3 thumbnail_extractor.py recording.mp4 ./thumbs 6

# YouTube video (download first)
yt-dlp -o "video.mp4" "https://youtube.com/watch?v=..."
python3 thumbnail_extractor.py video.mp4
Show full SKILL.md (378 more words)Show less
Output Files

For each candidate, the pipeline generates:

FileFormatDescription
{name}_{n}_{emotion}_{timestamp}_full.jpgJPG 95%Full video frame
{name}_{n}_{emotion}_{timestamp}_face.jpgJPG 95%Cropped face with padding
{name}_{n}_{emotion}_{timestamp}_transparent.pngPNG w/ alphaBackground-removed face cutout
{name}_manifest.jsonJSONMetadata for all candidates

Naming example: GMT20260130-210038_3_happy_2-10_full.jpg

  • GMT20260130-210038 — video name (truncated for Zoom recordings)
  • 3 — candidate number (ranked by score)
  • happy — detected dominant emotion
  • 2-10 — timestamp (2 minutes 10 seconds)
  • full / face / transparent — file type
Manifest JSON Structure
json
{
  "video": "GMT20260130-210038",
  "candidates": [
    {
      "index": 1,
      "timestamp": "2:10",
      "timestamp_sec": 130.0,
      "emotion": "happy",
      "emotion_score": 0.85,
      "combined_score": 12.4,
      "num_faces": 3,
      "is_presentation": false,
      "files": {
        "full": "..._full.jpg",
        "face_crop": "..._face.jpg",
        "transparent": "..._transparent.png"
      }
    }
  ]
}

Background Removal (Separate Step)

Since the Cowork sandbox may block model downloads, run rembg on the host Mac:

bash
# On host Mac (via osascript or Terminal)
cd ~/Downloads/thumb_candidates
python3 -c "
from rembg import remove
from PIL import Image
import glob, os

for f in sorted(glob.glob('*_face.jpg')):
    out = f.replace('_face.jpg', '_transparent.png')
    print(f'Processing {f}...')
    img = Image.open(f)
    result = remove(img)
    result.save(out)
    print(f'  -> {out} ({os.path.getsize(out)//1024}KB)')
"

This takes ~10-15 seconds per image on Apple Silicon. The u2net model downloads automatically on first run (~176MB).

Integration with Other Skills

Feeding into youtube-thumbnails

After extraction, use the transparent PNGs as compositing elements:

  1. Pick the best face cutout from the candidates
  2. Use it in the Gemini thumbnail prompt as a reference, or
  3. Composite it manually onto the generated Gemini background using ImageMagick:
bash
# Composite transparent face onto Gemini-generated background
convert gemini_background.jpg transparent_face.png \
  -gravity southeast -geometry +50+50 \
  -composite final_thumbnail.jpg
Pipeline flow
[Video MP4] → thumbnail-extraction → [face crops + transparent PNGs]
                                          ↓
                                   youtube-thumbnails → [Gemini background]
                                          ↓
                                   [Composite final thumbnail]

Tuning Parameters

Edit these at the top of thumbnail_extractor.py:

ParameterDefaultEffect
SAMPLE_INTERVAL_SEC10Lower = more frames scanned, slower
ANALYSIS_SCALE0.5Lower = faster face detection, less accurate
SCENE_THRESHOLD27.0Lower = more scene boundaries detected
MIN_FACE_CONFIDENCE0.80Higher = fewer false positive faces
top_n4Number of final candidates

For short videos (<10 min), consider SAMPLE_INTERVAL_SEC=5 for finer coverage.

Troubleshooting

  • OOM / killed process: The v2 pipeline never holds more than 1 frame in memory during Pass 1. If still OOM, increase SAMPLE_INTERVAL_SEC to 15-20.
  • All emotions "neutral": DeepFace model couldn't download (proxy block). Pass 1 smile detection still works — look at the num_smiles field in the manifest.
  • Face crop is wrong person: The pipeline picks the largest detected face. In screenshare mode, this may be a profile picture rather than a webcam face. Check the full frame to verify.
  • No faces detected: Zoom gallery recordings with "shared screen with gallery view" work best. Solo speaker view may have the face too close/large for the cascade detector — try lowering ANALYSIS_SCALE to 0.3.
  • Background removal artifacts: rembg's u2net can produce halos around hair. For cleaner results, try the u2net_human_seg model: remove(img, model_name='u2net_human_seg').
  • Slow processing: A 76-minute video takes ~2 minutes for Pass 1, ~15 seconds for Pass 2 (12 candidates), and ~60 seconds for bg removal (4 faces). Most time is in Pass 1 scanning.

© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in thumbnail-extraction of swyxio/skills.

  • SKILL.md
  • README.md
  • thumbnail_extractor.py

Open the folder on GitHubat commit 038ef34

Compare with similar skills

Thumbnail Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Thumbnail Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Thumbnail Extraction this skillswyxio/skills175—~2.4kAutomated safety check: PassMIT
Epub2podcast Gpt Imagedracohu2025-cloud/draco-skills-collection227—~1.9kAutomated safety check: PassMIT
Threads Carouselitchernetski/threads-carousel-claude-skill108—~3.5kAutomated safety check: PassMIT
Youtube Notetakersickn33/agentic-awesome-skills47k1 repos~2.3kAutomated safety check: PassMIT
Hook Writersocial-media-skills/skills128—~1.7kAutomated safety check: PassMIT
Nlm Skilliusztinpaul/ai-research-os-workshop1791 repos~6.9kAutomated safety check: PassMIT

Similar skills

  • Epub2podcast Gpt Image

    dracohu2025-cloud/draco-skills-collection

    可独立运行的 GPT-Image 增强版 EPUB2Podcast:在本地把 EPUB 转成双人中文音频、GPT-Image/Smart Slide 视觉页、最终 MP4,并生成 YouTube 发布素材。

    227 GitHub stars~1.9k tokensUpdated 22 days ago
    Media & CreativeAuto-check passed
  • Threads Carousel

    itchernetski/threads-carousel-claude-skill

    Convert text posts into visual carousel images or presentations for Threads, Instagram, LinkedIn, TikTok, YouTube.

    108 GitHub stars~3.5k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check passed
  • Youtube Notetaker

    sickn33/agentic-awesome-skills

    Turn YouTube talks into local study notes with slides, transcripts, editable annotations, and a markdown-backed viewer.

    47k GitHub starsUsed in 1 repo~2.3k tokens
    Documents & OfficeAuto-check passed
  • Hook Writer

    social-media-skills/skills

    A skill your agent uses to write the hook — the opening that earns attention — for any social content: a caption's first line, a video's first three seconds, a carousel cover slide, a thread opener…

    128 GitHub stars~1.7k tokensUpdated 8 days ago
    Media & CreativeAuto-check passed
  • Nlm Skill

    iusztinpaul/ai-research-os-workshop

    Expert guide for the NotebookLM CLI (nlm) and MCP server - interfaces for Google NotebookLM.

    179 GitHub starsUsed in 1 repo~6.9k tokens
    Knowledge ManagementAuto-check passed
  • Brand and Design Toolkit

    nextlevelbuilder/ui-ux-pro-max-skill

    Bundles design tasks behind one skill: brand identity, tokens, UI styling, logos, corporate identity mockups, slides, banners, icons and social images.

    134k GitHub starsUsed in 1 repo~3.5k tokens
    Media & CreativeAuto-check passed

More from swyxio/skills

All 89 skills in this repo
  • Programmatic Agents

    swyxio/skills

    Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Design, implement, audit, or refresh protected username and handle namespaces for public products.

    175 GitHub stars~1.1k tokensUpdated 4 days ago
    Auto-check passed
  • New Mac Setup

    swyxio/skills

    Fully automated new Mac setup for fullstack web developers and AI engineers.

    175 GitHub stars~4.3k tokensUpdated 4 days ago
    Auto-check passed
  • Youtube API

    swyxio/skills

    Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…

    175 GitHub stars~2.2k tokensUpdated 4 days ago
    Auto-check passed
  • Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.

    175 GitHub stars~1.5k tokensUpdated 4 days ago
    Auto-check: warnings
  • Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.

    175 GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed

Works with

Questions about Thumbnail Extraction

What does Thumbnail Extraction do?

Extracts the most interesting frames from video files for thumbnail compositing. Thumbnail Extraction is an agent skill from swyxio/skills. Extracts the most interesting frames from video files for thumbnail compositing.

When should I use Thumbnail Extraction?

Thumbnail Extraction fits situations like: asked to extract thumbnails; find interesting frames; grab screenshots from video; create thumbnail candidates from recordings.

How do I install Thumbnail Extraction in Claude Code?

Run `npx skills add swyxio/skills --skill thumbnail-extraction -a claude-code`. Or copy the skill folder (thumbnail-extraction in swyxio/skills) into .claude/skills/thumbnail-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Thumbnail Extraction in Codex?

Run `npx skills add swyxio/skills --skill thumbnail-extraction -a codex`. Or copy the skill folder (thumbnail-extraction in swyxio/skills) into .agents/skills/thumbnail-extraction in your project. Codex loads it when a task matches its description.

Can I use Thumbnail Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill thumbnail-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/thumbnail-extraction, .gemini/skills/thumbnail-extraction, .github/skills/thumbnail-extraction and .opencode/skills/thumbnail-extraction in your project.

What does Thumbnail Extraction need to run?

Going by SKILL.md and its folder, Thumbnail Extraction needs Python for the scripts in its folder and the command-line tools its instructions call (python3, pip, pip3 and yt-dlp). Our summary lists: Python 3.

Does Thumbnail Extraction access the network?

SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Thumbnail Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Thumbnail Extraction use?

Thumbnail Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Thumbnail Extraction use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Thumbnail Extraction?

Skills that share tags, products or a category with Thumbnail Extraction: Epub2podcast Gpt Image (dracohu2025-cloud/draco-skills-collection, 227 stars), Threads Carousel (itchernetski/threads-carousel-claude-skill, 108 stars), Youtube Notetaker (sickn33/agentic-awesome-skills, 47k stars) and Hook Writer (social-media-skills/skills, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Thumbnail Extraction?

swyxio (a GitHub user) maintains it in swyxio/skills, which has 175 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.

Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.