Agent skill

Video Perception

by jordanrendric in jordanrendric/claude-video-vision

A skill your agent uses when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation

MITAuto-check passedMedia & Creative

Install Video Perception

skills CLI
$ npx skills add jordanrendric/claude-video-vision --skill video-perception -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jordanrendric/claude-video-vision video-perception --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jordanrendric/claude-video-vision.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/video-perception .claude/skills/video-perception && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-perception
GitHub stars
1.4k
Token cost
~1.4k tokens
SKILL.md length
728 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation

  • Works in 6 steps: Always start with video_info to get… → REQUIRED for videos > 30s: Call… → Use the analysis results and… → …
  • The user mentions a video file (.mp4
  • SKILL.md covers Available Tools, Workflow, Parameter Guide and Working with Results
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Video Perception is an agent skill from jordanrendric/claude-video-vision. Use when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative. It works with YouTube, Model Context Protocol, FFmpeg and Google Gemini. The repository describes itself as: Give Claude the ability to watch and understand videos — Claude Code plugin with frame extraction and multimodal audio analysis. The licence is MIT.

When your agent uses it

  • The user mentions a video file (.mp4
  • Asks to watch/analyze/review a video
  • References video content in conversation

Example prompts

  • “/video-perception”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Always start with video_info to get duration, resolution, and audio presence.
  2. REQUIRED for videos > 30s: Call video_analyze BEFORE extracting any frames.
  3. Use the analysis results and transcription to plan your frame extraction strategy
  4. Call video_watch to extract frames
  5. Use video_detail to drill into specific moments
  6. When the user asks follow-up questions about the same video, consult

What it can do on your machine

Read from SKILL.md and the folder at commit 4a4f990. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Perception loads about 1.4k tokens when it runs. Until then it costs about 47 tokens; SKILL.md has 728 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~47
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jordanrendric/claude-video-vision at commit 4a4f990, republished under its MIT licence (© jordanrendric). 728 words, ~1,400 tokens.

Download SKILL.mdSave it as .claude/skills/video-perception/SKILL.md (or your agent's skills folder).
name
video-perception
description
Use when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation

Video Perception

You have access to video understanding tools via the claude-video-vision MCP server.

Available Tools

  • video_analyze — Analyze video structure with ffmpeg filters (scene changes, silence, motion, etc.). Use this BEFORE extracting frames to plan your strategy.
  • video_watch — Extract frames + process audio from a video. Supports variable FPS/resolution per segment.
  • video_detail — Drill into specific segments. Separates extraction from viewing — extract many frames, view few at a time.
  • video_info — Get video metadata without processing.
  • video_configure — Change settings (backend, resolution, enable_index, etc.).
  • video_setup — Check/install dependencies.

Workflow

IMPORTANT: You MUST follow these steps in order. Do NOT skip step 2.

  1. Always start with video_info to get duration, resolution, and audio presence. If the user gives a YouTube URL, pass the URL directly as path. The MCP server downloads it with yt-dlp, prefers YouTube subtitles/auto-captions for transcription, and falls back to the configured audio backend only when captions are missing, empty, or suspiciously incomplete.

  2. REQUIRED for videos > 30s: Call video_analyze BEFORE extracting any frames. This is NOT optional — it gives you structural data to make smart extraction decisions. Select filters relevant to the user's question:

    User intentFilters to select
    "What happens in this video?"scene_changes, silence, transcription
    "Find the scene transitions"scene_changes, black_intervals
    "Are there frozen/stuck parts?"freeze, blur
    "Is this a talking head or action?"motion
    "When does the music start?"silence, loudness
    "Analyze the lighting"exposure
    "Summarize this lecture"transcription, scene_changes, silence
    General / unclear intentscene_changes, silence, transcription

    Always include transcription: true when the video has audio — the transcription tells you WHERE to look visually.

    scene_changes: true reports hard cuts (scdet score >= 8). If the user needs softer transitions (dissolves, slow fades), pass scene_changes: { threshold: 4 }; if handheld or fast-moving footage floods the list, raise it (e.g. { threshold: 15 }).

  3. Use the analysis results and transcription to plan your frame extraction strategy:

    • Low FPS (0.1-0.5) for static or predictable segments
    • Higher FPS (1-3) only around scene changes, motion peaks, or moments referenced in speech ("look at this", "as you can see", "let me show you")
    • Never exceed the minimum FPS needed for the task
    • Prefer fewer segments at lower FPS — you can always drill deeper
  4. Call video_watch to extract frames:

    • For short videos (< 2 minutes): Use fps: "auto" without view_sample — short videos need full coverage to avoid missing brief moments. The auto FPS already adapts to duration.
    • For long videos (> 2 minutes): Use segments based on analysis data with variable FPS, and view_sample to limit initial frame count. You can always drill deeper with video_detail.
  5. Use video_detail to drill into specific moments:

    • Start with 3-5 second windows around points of interest
    • Use view_sample: 3 to preview (first, middle, last frame)
    • Then request specific timestamps with view if you need more detail
    • Expand the window only if the initial view is insufficient
    • Treat frame viewing like a binary search — narrow down to what matters
    • Never view all extracted frames at once
  6. When the user asks follow-up questions about the same video, consult the manifest already in your context. Do not re-extract frames you already have at the same resolution. Do not re-request frames you already have in context.

Show full SKILL.md (207 more words)Show less

Parameter Guide

fps: "auto" for general overview. Use the video's original fps (from video_info) for frame-by-frame detail. Use 5-10 for analyzing specific short moments. Use 0.1-0.5 for long videos.

resolution: 256-512 for quick scans. 512-768 for normal analysis. 1024+ when reading on-screen text or fine details.

segments: Use when you have analysis data. Each segment can have its own fps and resolution. Overrides global fps/start_time/end_time.

view_sample: Returns N evenly spaced frames from the extracted set. Use this to avoid flooding context with too many images.

skip_audio: Set to true when you only need visual analysis.

YouTube URLs: Pass supported YouTube URLs directly as path. Treat transcription_source: "youtube_subtitles" as stronger than youtube_auto_captions; auto-captions can still have recognition errors.

Working with Results

You receive:

  • Manifest (when enable_index is on) — index of all cached frames by resolution and timestamp. Use this to avoid redundant requests.
  • Frames as images — look at them to understand what's happening visually
  • Audio transcription with timestamps — read the speech content
  • Audio tags — non-speech events (music, sounds, etc.)
  • Analysis data — scene changes, silence intervals, motion levels, etc.

Combine all sources to form a complete understanding. Use analysis + transcription to guide where you look visually. The analysis tells you WHEN things happen; the frames tell you WHAT happens.

© jordanrendric, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/video-perception of jordanrendric/claude-video-vision.

Open the folder on GitHubat commit 4a4f990

Compare with similar skills

Video Perception next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Perception compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Perception this skilljordanrendric/claude-video-vision1.4k—~1.4kAutomated safety check: PassMIT
Watch Videocoreyhaines31/makerskills850—~3.8kAutomated safety check: PassMIT
Youtubeeat-pray-ai/yutu699—~1.1kAutomated safety check: PassMIT
Youtube Clipperop7418/Youtube-clipper-skill2.2k—~1.6kAutomated safety check: NotesMIT
Video Transcribewendy7756/AI-Video-Transcriber3.3k—~937Automated safety check: NotesApache-2.0
Gbro Collage Brollpyang5166/gbro-collage-broll1.3k—~2.6kAutomated safety check: NotesMIT

Similar skills

  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    850 GitHub stars~3.8k tokensUpdated today
    Media & CreativeAuto-check passed
  • Youtube

    eat-pray-ai/yutu

    A skill your agent uses whenever the user mentions YouTube, video uploads, channel management, playlists, video SEO, or any YouTube Data API operation.

    699 GitHub stars~1.1k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Youtube Clipper

    op7418/Youtube-clipper-skill

    YouTube 视频智能剪辑工具。下载视频和字幕,AI 分析生成精细章节(几分钟级别), 用户选择片段后自动剪辑、翻译字幕为中英双语、烧录字幕到视频,并生成总结文案。

    2.2k GitHub stars~1.6k tokensUpdated 8 mo ago
    Media & CreativeAuto-check: notes
  • Video Transcribe

    wendy7756/AI-Video-Transcriber

    Transcribe and summarize a video or podcast from a URL (YouTube, TikTok, Bilibili, Apple Podcasts, SoundCloud, 30+ platforms) or from a local media/.txt file.

    3.3k GitHub stars~937 tokensUpdated 24 days ago
    Media & CreativeAuto-check: notes
  • Gbro Collage Broll

    pyang5166/gbro-collage-broll

    将约 5 秒口播文稿、观点句或抽象概念做成高级 editorial halftone paper-collage / 半调纸拼贴 B-roll。用户说“collage b-roll”“纸拼贴 b-roll”“半调拼贴”“拼贴风格配画面”“用这段文稿做拼贴动画”“gbro-collage-broll”,或希望把一句文稿转成拼贴视觉隐喻时,必须使用此…

    1.3k GitHub stars~2.6k tokensUpdated 2 mo ago
    Media & CreativeAuto-check: notes
  • Douyin Video

    yzfly/douyin-mcp-server

    抖音无水印视频下载和文案提取工具. An agent skill from yzfly/douyin-mcp-server.

    1.3k GitHub stars~666 tokensUpdated 3 mo ago
    Media & CreativeAuto-check passed

Questions about Video Perception

What does Video Perception do?

A skill your agent uses when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation. Video Perception is an agent skill from jordanrendric/claude-video-vision.

When should I use Video Perception?

Video Perception fits situations like: the user mentions a video file (.mp4; asks to watch/analyze/review a video; references video content in conversation.

How do I install Video Perception in Claude Code?

Run `npx skills add jordanrendric/claude-video-vision --skill video-perception -a claude-code`. Or copy the skill folder (skills/video-perception in jordanrendric/claude-video-vision) into .claude/skills/video-perception in your project. Claude Code loads it when a task matches its description.

How do I install Video Perception in Codex?

Run `npx skills add jordanrendric/claude-video-vision --skill video-perception -a codex`. Or copy the skill folder (skills/video-perception in jordanrendric/claude-video-vision) into .agents/skills/video-perception in your project. Codex loads it when a task matches its description.

Can I use Video Perception in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jordanrendric/claude-video-vision --skill video-perception -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-perception, .gemini/skills/video-perception, .github/skills/video-perception and .opencode/skills/video-perception in your project.

What does Video Perception need to run?

SKILL.md names no scripts, command-line tools or credentials: Video Perception is instructions for the agent only.

Does Video Perception access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Video Perception safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Video Perception use?

Video Perception is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Perception use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Perception?

Skills that share tags, products or a category with Video Perception: Watch Video (coreyhaines31/makerskills, 850 stars), Youtube (eat-pray-ai/yutu, 699 stars), Youtube Clipper (op7418/Youtube-clipper-skill, 2.2k stars) and Video Transcribe (wendy7756/AI-Video-Transcriber, 3.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Perception?

jordanrendric (a GitHub user) maintains it in jordanrendric/claude-video-vision, which has 1,354 GitHub stars. The repository was last updated on October 7, 2026.

Source: jordanrendric/claude-video-vision on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.