Epub2podcast Gpt Image
dracohu2025-cloud/draco-skills-collection
可独立运行的 GPT-Image 增强版 EPUB2Podcast:在本地把 EPUB 转成双人中文音频、GPT-Image/Smart Slide 视觉页、最终 MP4,并生成 YouTube 发布素材。
Extracts the most interesting frames from video files for thumbnail compositing.
$ npx skills add swyxio/skills --skill thumbnail-extraction -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install swyxio/skills thumbnail-extraction --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/thumbnail-extraction .claude/skills/thumbnail-extraction && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "thumbnail-extraction" agent skill from https://github.com/swyxio/skills/tree/main/thumbnail-extraction into .claude/skills/thumbnail-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thumbnail-extraction", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/swyxio/skills/tree/main/thumbnail-extractionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add swyxio/skills --skill thumbnail-extraction -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install swyxio/skills thumbnail-extraction --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/thumbnail-extraction .agents/skills/thumbnail-extraction && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "thumbnail-extraction" agent skill from https://github.com/swyxio/skills/tree/main/thumbnail-extraction into .agents/skills/thumbnail-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thumbnail-extraction", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill thumbnail-extraction -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install swyxio/skills thumbnail-extraction --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/thumbnail-extraction .cursor/skills/thumbnail-extraction && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "thumbnail-extraction" agent skill from https://github.com/swyxio/skills/tree/main/thumbnail-extraction into .cursor/skills/thumbnail-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thumbnail-extraction", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/swyxio/skills.git --path thumbnail-extraction--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add swyxio/skills --skill thumbnail-extraction -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install swyxio/skills thumbnail-extraction --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/thumbnail-extraction .gemini/skills/thumbnail-extraction && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "thumbnail-extraction" agent skill from https://github.com/swyxio/skills/tree/main/thumbnail-extraction into .gemini/skills/thumbnail-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thumbnail-extraction", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install swyxio/skills thumbnail-extractionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add swyxio/skills --skill thumbnail-extraction -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/thumbnail-extraction .github/skills/thumbnail-extraction && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "thumbnail-extraction" agent skill from https://github.com/swyxio/skills/tree/main/thumbnail-extraction into .github/skills/thumbnail-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thumbnail-extraction", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add swyxio/skills --skill thumbnail-extraction -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install swyxio/skills thumbnail-extraction --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/swyxio/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/thumbnail-extraction .opencode/skills/thumbnail-extraction && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "thumbnail-extraction" agent skill from https://github.com/swyxio/skills/tree/main/thumbnail-extraction into .opencode/skills/thumbnail-extraction/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "thumbnail-extraction", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
thumbnail-extractionExtracts the most interesting frames from video files for thumbnail compositing.
Thumbnail Extraction is an agent skill from swyxio/skills. Extracts the most interesting frames from video files for thumbnail compositing. Detects faces, expressions, smiles, and presentation slides. Outputs full frames, face crops, and transparent cutouts. Use when asked to extract thumbnails, find interesting frames, grab screenshots from video, or create thumbnail candidates from recordings.
Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `README.md` and `thumbnail_extractor.py`).
It sits in Media & Creative, covering Drug discovery and cheminformatics and Slides and decks. It works with YouTube. The repository describes itself as: Agent skills for Claude Code and other AI agents. The licence is MIT.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 038ef34. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
python3pippip3yt-dlpFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
youtube.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Thumbnail Extraction loads about 2.4k tokens when it runs. Until then it costs about 90 tokens; SKILL.md has 917 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from swyxio/skills at commit 038ef34, republished under its MIT licence (© swyxio). 917 words, ~2,392 tokens.
.claude/skills/thumbnail-extraction/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Automatically scan a local MP4 video (or YouTube URL via yt-dlp) and extract the 4 most visually interesting frames — prioritizing expressive faces (laughing, shocked, smiling) and engaging presentation slides. Outputs full frames, face crops, and background-removed transparent PNGs ready for compositing.
youtube-thumbnails skill)# In sandbox (Cowork VM):
pip install opencv-python scenedetect deepface pillow numpy --break-system-packages
# On host Mac (for background removal — sandbox can't download the model):
pip3 install 'rembg[cpu]' pillow --break-system-packagesffmpeg (usually pre-installed)python3 (3.10+)yt-dlp (optional, for YouTube URLs): pip install yt-dlp --break-system-packagesPass 1 — Quick Scan (OpenCV only, no deep learning)
Pass 2 — Deep Analysis (DeepFace, only on top 12 candidates)
Pass 3 — Output (rembg, only on final 4 frames)
| Signal | Weight | Notes |
|---|---|---|
| Face detected | +2.0 per face (cap 3) | Gallery views score high |
| Smile detected | +3.0 per smile | Cascade-based, no model needed |
| Smile size ratio | +5.0 × ratio | Bigger smiles = more expressive |
| Multi-person shot | +1.0 bonus | 2+ faces = engaging |
| Happy expression | +2.0 bonus (Pass 2) | Best for thumbnails |
| Surprise expression | +2.0 bonus (Pass 2) | Eye-catching |
| Fear/angry expression | +1.0 bonus (Pass 2) | "Shocked" reactions |
| Visual variance | +0.0–1.5 | Normalized by frame complexity |
| Presentation slide | baseline 1.5 | Useful for slide screenshots |
The pipeline enforces temporal spread to avoid clustering picks in one segment:
top_n equal time segments, pick best from eachThis ensures a 76-minute video yields picks from different parts (e.g., 1:00, 2:10, 21:50, 48:50) rather than clustering in the most face-heavy section.
python3 thumbnail_extractor.py <video_path> [output_dir] [top_n]Arguments:
video_path — Path to MP4 file (required)output_dir — Where to save outputs (default: ~/Downloads/thumb_candidates)top_n — Number of candidates to extract (default: 4)Examples:
# Basic — extract 4 best frames
python3 thumbnail_extractor.py "GMT20260130-210038_Recording_gallery_2380x1544.mp4"
# Custom output dir and count
python3 thumbnail_extractor.py recording.mp4 ./thumbs 6
# YouTube video (download first)
yt-dlp -o "video.mp4" "https://youtube.com/watch?v=..."
python3 thumbnail_extractor.py video.mp4For each candidate, the pipeline generates:
| File | Format | Description |
|---|---|---|
{name}_{n}_{emotion}_{timestamp}_full.jpg | JPG 95% | Full video frame |
{name}_{n}_{emotion}_{timestamp}_face.jpg | JPG 95% | Cropped face with padding |
{name}_{n}_{emotion}_{timestamp}_transparent.png | PNG w/ alpha | Background-removed face cutout |
{name}_manifest.json | JSON | Metadata for all candidates |
Naming example: GMT20260130-210038_3_happy_2-10_full.jpg
GMT20260130-210038 — video name (truncated for Zoom recordings)3 — candidate number (ranked by score)happy — detected dominant emotion2-10 — timestamp (2 minutes 10 seconds)full / face / transparent — file type{
"video": "GMT20260130-210038",
"candidates": [
{
"index": 1,
"timestamp": "2:10",
"timestamp_sec": 130.0,
"emotion": "happy",
"emotion_score": 0.85,
"combined_score": 12.4,
"num_faces": 3,
"is_presentation": false,
"files": {
"full": "..._full.jpg",
"face_crop": "..._face.jpg",
"transparent": "..._transparent.png"
}
}
]
}Since the Cowork sandbox may block model downloads, run rembg on the host Mac:
# On host Mac (via osascript or Terminal)
cd ~/Downloads/thumb_candidates
python3 -c "
from rembg import remove
from PIL import Image
import glob, os
for f in sorted(glob.glob('*_face.jpg')):
out = f.replace('_face.jpg', '_transparent.png')
print(f'Processing {f}...')
img = Image.open(f)
result = remove(img)
result.save(out)
print(f' -> {out} ({os.path.getsize(out)//1024}KB)')
"This takes ~10-15 seconds per image on Apple Silicon. The u2net model downloads automatically on first run (~176MB).
youtube-thumbnailsAfter extraction, use the transparent PNGs as compositing elements:
# Composite transparent face onto Gemini-generated background
convert gemini_background.jpg transparent_face.png \
-gravity southeast -geometry +50+50 \
-composite final_thumbnail.jpg[Video MP4] → thumbnail-extraction → [face crops + transparent PNGs]
↓
youtube-thumbnails → [Gemini background]
↓
[Composite final thumbnail]Edit these at the top of thumbnail_extractor.py:
| Parameter | Default | Effect |
|---|---|---|
SAMPLE_INTERVAL_SEC | 10 | Lower = more frames scanned, slower |
ANALYSIS_SCALE | 0.5 | Lower = faster face detection, less accurate |
SCENE_THRESHOLD | 27.0 | Lower = more scene boundaries detected |
MIN_FACE_CONFIDENCE | 0.80 | Higher = fewer false positive faces |
top_n | 4 | Number of final candidates |
For short videos (<10 min), consider SAMPLE_INTERVAL_SEC=5 for finer coverage.
SAMPLE_INTERVAL_SEC to 15-20.num_smiles field in the manifest.ANALYSIS_SCALE to 0.3.u2net_human_seg model: remove(img, model_name='u2net_human_seg').© swyxio, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in thumbnail-extraction of swyxio/skills.
Open the folder on GitHubat commit 038ef34
Thumbnail Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Thumbnail Extraction this skillswyxio/skills | 175 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Epub2podcast Gpt Imagedracohu2025-cloud/draco-skills-collection | 227 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Threads Carouselitchernetski/threads-carousel-claude-skill | 108 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Youtube Notetakersickn33/agentic-awesome-skills | 47k | 1 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Hook Writersocial-media-skills/skills | 128 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Nlm Skilliusztinpaul/ai-research-os-workshop | 179 | 1 repos | ~6.9k | Automated safety check: Pass | MIT |
dracohu2025-cloud/draco-skills-collection
可独立运行的 GPT-Image 增强版 EPUB2Podcast:在本地把 EPUB 转成双人中文音频、GPT-Image/Smart Slide 视觉页、最终 MP4,并生成 YouTube 发布素材。
itchernetski/threads-carousel-claude-skill
Convert text posts into visual carousel images or presentations for Threads, Instagram, LinkedIn, TikTok, YouTube.
sickn33/agentic-awesome-skills
Turn YouTube talks into local study notes with slides, transcripts, editable annotations, and a markdown-backed viewer.
social-media-skills/skills
A skill your agent uses to write the hook — the opening that earns attention — for any social content: a caption's first line, a video's first three seconds, a carousel cover slide, a thread opener…
iusztinpaul/ai-research-os-workshop
Expert guide for the NotebookLM CLI (nlm) and MCP server - interfaces for Google NotebookLM.
nextlevelbuilder/ui-ux-pro-max-skill
Bundles design tasks behind one skill: brand identity, tokens, UI styling, logos, corporate identity mockups, slides, banners, icons and social images.
swyxio/skills
Run a selected coding-agent CLI programmatically, with latency, error, usage, cost, and trace logging.
swyxio/skills
Design, implement, audit, or refresh protected username and handle namespaces for public products.
swyxio/skills
Fully automated new Mac setup for fullstack web developers and AI engineers.
swyxio/skills
Manage YouTube videos programmatically via the YouTube Data API v3 — upload video files, upload custom thumbnails, update video metadata (titles, descriptions, tags), and query video/channel info…
swyxio/skills
Batch YouTube Studio upload workflow for videos sourced from Airtable, Google Drive, Loom, YouTube, or local files.
swyxio/skills
Reconstruct and visually analyze paired agent, game, or policy trajectories to determine whether changed actions produced their intended effects.
Works with
Categories
Extracts the most interesting frames from video files for thumbnail compositing. Thumbnail Extraction is an agent skill from swyxio/skills. Extracts the most interesting frames from video files for thumbnail compositing.
Thumbnail Extraction fits situations like: asked to extract thumbnails; find interesting frames; grab screenshots from video; create thumbnail candidates from recordings.
Run `npx skills add swyxio/skills --skill thumbnail-extraction -a claude-code`. Or copy the skill folder (thumbnail-extraction in swyxio/skills) into .claude/skills/thumbnail-extraction in your project. Claude Code loads it when a task matches its description.
Run `npx skills add swyxio/skills --skill thumbnail-extraction -a codex`. Or copy the skill folder (thumbnail-extraction in swyxio/skills) into .agents/skills/thumbnail-extraction in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add swyxio/skills --skill thumbnail-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/thumbnail-extraction, .gemini/skills/thumbnail-extraction, .github/skills/thumbnail-extraction and .opencode/skills/thumbnail-extraction in your project.
Going by SKILL.md and its folder, Thumbnail Extraction needs Python for the scripts in its folder and the command-line tools its instructions call (python3, pip, pip3 and yt-dlp). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Thumbnail Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Thumbnail Extraction: Epub2podcast Gpt Image (dracohu2025-cloud/draco-skills-collection, 227 stars), Threads Carousel (itchernetski/threads-carousel-claude-skill, 108 stars), Youtube Notetaker (sickn33/agentic-awesome-skills, 47k stars) and Hook Writer (social-media-skills/skills, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
swyxio (a GitHub user) maintains it in swyxio/skills, which has 175 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 5, 2026.
Source: swyxio/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.