Watch Video
coreyhaines31/makerskills
When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.
A skill your agent uses when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation
$ npx skills add jordanrendric/claude-video-vision --skill video-perception -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jordanrendric/claude-video-vision video-perception --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jordanrendric/claude-video-vision.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-perception .claude/skills/video-perception && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "video-perception" agent skill from https://github.com/jordanrendric/claude-video-vision/tree/main/skills/video-perception into .claude/skills/video-perception/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-perception", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jordanrendric/claude-video-vision/tree/main/skills/video-perceptionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jordanrendric/claude-video-vision --skill video-perception -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jordanrendric/claude-video-vision video-perception --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jordanrendric/claude-video-vision.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/video-perception .agents/skills/video-perception && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "video-perception" agent skill from https://github.com/jordanrendric/claude-video-vision/tree/main/skills/video-perception into .agents/skills/video-perception/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-perception", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jordanrendric/claude-video-vision --skill video-perception -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jordanrendric/claude-video-vision video-perception --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jordanrendric/claude-video-vision.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/video-perception .cursor/skills/video-perception && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "video-perception" agent skill from https://github.com/jordanrendric/claude-video-vision/tree/main/skills/video-perception into .cursor/skills/video-perception/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-perception", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jordanrendric/claude-video-vision.git --path skills/video-perception--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jordanrendric/claude-video-vision --skill video-perception -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jordanrendric/claude-video-vision video-perception --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jordanrendric/claude-video-vision.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/video-perception .gemini/skills/video-perception && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "video-perception" agent skill from https://github.com/jordanrendric/claude-video-vision/tree/main/skills/video-perception into .gemini/skills/video-perception/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-perception", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jordanrendric/claude-video-vision video-perceptionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jordanrendric/claude-video-vision --skill video-perception -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jordanrendric/claude-video-vision.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/video-perception .github/skills/video-perception && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "video-perception" agent skill from https://github.com/jordanrendric/claude-video-vision/tree/main/skills/video-perception into .github/skills/video-perception/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-perception", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jordanrendric/claude-video-vision --skill video-perception -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jordanrendric/claude-video-vision video-perception --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jordanrendric/claude-video-vision.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/video-perception .opencode/skills/video-perception && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "video-perception" agent skill from https://github.com/jordanrendric/claude-video-vision/tree/main/skills/video-perception into .opencode/skills/video-perception/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "video-perception", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
video-perceptionA skill your agent uses when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation
Video Perception is an agent skill from jordanrendric/claude-video-vision. Use when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Media & Creative. It works with YouTube, Model Context Protocol, FFmpeg and Google Gemini. The repository describes itself as: Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis. The licence is MIT.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4a4f990. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Video Perception loads about 1.4k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 728 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jordanrendric/claude-video-vision at commit 4a4f990, republished under its MIT licence (© jordanrendric). 728 words, ~1,400 tokens.
.claude/skills/video-perception/SKILL.md (or your agent's skills folder).You have access to video understanding tools via the claude-video-vision MCP server.
video_analyze — Analyze video structure with ffmpeg filters (scene changes, silence, motion, etc.). Use this BEFORE extracting frames to plan your strategy.video_watch — Extract frames + process audio from a video. Supports variable FPS/resolution per segment.video_detail — Drill into specific segments. Separates extraction from viewing — extract many frames, view few at a time.video_info — Get video metadata without processing.video_configure — Change settings (backend, resolution, enable_index, etc.).video_setup — Check/install dependencies.IMPORTANT: You MUST follow these steps in order. Do NOT skip step 2.
Always start with video_info to get duration, resolution, and audio presence.
If the user gives a YouTube URL, pass the URL directly as path.
The MCP server downloads it with yt-dlp, prefers YouTube subtitles/auto-captions
for transcription, and falls back to the configured audio backend only when
captions are missing, empty, or suspiciously incomplete.
REQUIRED for videos > 30s: Call video_analyze BEFORE extracting any frames.
This is NOT optional — it gives you structural data to make smart extraction decisions.
Select filters relevant to the user's question:
| User intent | Filters to select |
|---|---|
| "What happens in this video?" | scene_changes, silence, transcription |
| "Find the scene transitions" | scene_changes, black_intervals |
| "Are there frozen/stuck parts?" | freeze, blur |
| "Is this a talking head or action?" | motion |
| "When does the music start?" | silence, loudness |
| "Analyze the lighting" | exposure |
| "Summarize this lecture" | transcription, scene_changes, silence |
| General / unclear intent | scene_changes, silence, transcription |
Always include transcription: true when the video has audio — the transcription
tells you WHERE to look visually.
scene_changes: true reports hard cuts (scdet score >= 8). If the user needs softer
transitions (dissolves, slow fades), pass scene_changes: { threshold: 4 }; if
handheld or fast-moving footage floods the list, raise it (e.g. { threshold: 15 }).
Use the analysis results and transcription to plan your frame extraction strategy:
Call video_watch to extract frames:
fps: "auto" without view_sample — short videos need full coverage to avoid missing brief moments. The auto FPS already adapts to duration.segments based on analysis data with variable FPS, and view_sample to limit initial frame count. You can always drill deeper with video_detail.Use video_detail to drill into specific moments:
view_sample: 3 to preview (first, middle, last frame)view if you need more detailWhen the user asks follow-up questions about the same video, consult the manifest already in your context. Do not re-extract frames you already have at the same resolution. Do not re-request frames you already have in context.
fps: "auto" for general overview. Use the video's original fps (from video_info) for frame-by-frame detail. Use 5-10 for analyzing specific short moments. Use 0.1-0.5 for long videos.
resolution: 256-512 for quick scans. 512-768 for normal analysis. 1024+ when reading on-screen text or fine details.
segments: Use when you have analysis data. Each segment can have its own fps and resolution. Overrides global fps/start_time/end_time.
view_sample: Returns N evenly spaced frames from the extracted set. Use this to avoid flooding context with too many images.
skip_audio: Set to true when you only need visual analysis.
YouTube URLs: Pass supported YouTube URLs directly as path. Treat
transcription_source: "youtube_subtitles" as stronger than
youtube_auto_captions; auto-captions can still have recognition errors.
You receive:
Combine all sources to form a complete understanding. Use analysis + transcription to guide where you look visually. The analysis tells you WHEN things happen; the frames tell you WHAT happens.
© jordanrendric, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/video-perception of jordanrendric/claude-video-vision.
Open the folder on GitHubat commit 4a4f990
Video Perception next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Video Perception this skilljordanrendric/claude-video-vision | 1.4k | — | ~1.4k | Automated safety check: Pass | MIT | |
| Watch Videocoreyhaines31/makerskills | 850 | — | ~3.8k | Automated safety check: Pass | MIT | |
| Youtubeeat-pray-ai/yutu | 699 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Youtube Clipperop7418/Youtube-clipper-skill | 2.2k | — | ~1.6k | Automated safety check: Notes | MIT | |
| Video Transcribewendy7756/AI-Video-Transcriber | 3.3k | — | ~937 | Automated safety check: Notes | Apache-2.0 | |
| Gbro Collage Brollpyang5166/gbro-collage-broll | 1.3k | — | ~2.6k | Automated safety check: Notes | MIT |
coreyhaines31/makerskills
When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.
eat-pray-ai/yutu
A skill your agent uses whenever the user mentions YouTube, video uploads, channel management, playlists, video SEO, or any YouTube Data API operation.
op7418/Youtube-clipper-skill
YouTube 视频智能剪辑工具。下载视频和字幕,AI 分析生成精细章节(几分钟级别), 用户选择片段后自动剪辑、翻译字幕为中英双语、烧录字幕到视频,并生成总结文案。
wendy7756/AI-Video-Transcriber
Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.
pyang5166/gbro-collage-broll
将约 5 秒口播文稿、观点句或抽象概念做成高级 editorial halftone paper-collage / 半调纸拼贴 B-roll。用户说“collage b-roll”“纸拼贴 b-roll”“半调拼贴”“拼贴风格配画面”“用这段文稿做拼贴动画”“gbro-collage-broll”,或希望把一句文稿转成拼贴视觉隐喻时,必须使用此…
yzfly/douyin-mcp-server
抖音无水印视频下载和文案提取工具. An agent skill from yzfly/douyin-mcp-server.
Categories
A skill your agent uses when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation. Video Perception is an agent skill from jordanrendric/claude-video-vision.
Video Perception fits situations like: the user mentions a video file (.mp4; asks to watch/analyze/review a video; references video content in conversation.
Run `npx skills add jordanrendric/claude-video-vision --skill video-perception -a claude-code`. Or copy the skill folder (skills/video-perception in jordanrendric/claude-video-vision) into .claude/skills/video-perception in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jordanrendric/claude-video-vision --skill video-perception -a codex`. Or copy the skill folder (skills/video-perception in jordanrendric/claude-video-vision) into .agents/skills/video-perception in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jordanrendric/claude-video-vision --skill video-perception -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-perception, .gemini/skills/video-perception, .github/skills/video-perception and .opencode/skills/video-perception in your project.
SKILL.md names no scripts, command-line tools or credentials: Video Perception is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Video Perception is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Video Perception: Watch Video (coreyhaines31/makerskills, 850 stars), Youtube (eat-pray-ai/yutu, 699 stars), Youtube Clipper (op7418/Youtube-clipper-skill, 2.2k stars) and Video Transcribe (wendy7756/AI-Video-Transcriber, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jordanrendric (a GitHub user) maintains it in jordanrendric/claude-video-vision, which has 1,354 GitHub stars. The repository was last updated on October 7, 2026.
Source: jordanrendric/claude-video-vision on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.