Agent skill

Video Understand

by calesthio in calesthio/OpenMontage

Understand video content locally using ffmpeg frame extraction and Whisper transcription.

AGPL-3.0Auto-check passedMedia & Creative

Install Video Understand

skills CLI
$ npx skills add calesthio/OpenMontage --skill video-understand -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install calesthio/OpenMontage video-understand --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/calesthio/OpenMontage.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/video-understand .claude/skills/video-understand && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-understand
GitHub stars
66k
Token cost
~841 tokens
SKILL.md length
190 words
Files
3 (incl. scripts, references)
Skills in repo
41
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Understand video content locally using ffmpeg frame extraction and Whisper transcription.

  • Understanding what a video contains
  • SKILL.md covers Prerequisites, Commands, CLI Options and Extraction Modes, plus 2 more sections
  • Runs Python scripts from its folder; calls python3, brew and pip
  • Transcribing video audio locally

What it does

Video Understand is an agent skill from calesthio/OpenMontage. Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without API keys.

Its SKILL.md is about 840 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts and reference files (for example `references/output-format.md` and `scripts/understand_video.py`).

It sits in Media & Creative, covering Video production and Transcription. It works with FFmpeg. The repository describes itself as: World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant… The licence is AGPL-3.0.

When your agent uses it

  • Understanding what a video contains
  • Transcribing video audio locally
  • Extracting key frames for visual analysis
  • Getting video content without API keys

Example prompts

  • “/video-understand”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9327439. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • brew
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Understand loads about 841 tokens when it runs, and up to ~1.6k if it reads all its reference files. Until then it costs about 77 tokens; SKILL.md has 190 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~77
When it runs · the whole SKILL.md, loaded when a task matches
~841
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from calesthio/OpenMontage at commit 9327439, republished under its AGPL-3.0 licence (© calesthio). 190 words, ~841 tokens.

Download SKILL.mdSave it as .claude/skills/video-understand/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
video-understand
description
Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without API keys.

video-understand

Understand video content locally using ffmpeg for frame extraction and Whisper for transcription. Fully offline, no API keys required.

Prerequisites

  • ffmpeg + ffprobe (required): brew install ffmpeg
  • openai-whisper (optional, for transcription): pip install openai-whisper

Commands

bash
# Scene detection + transcribe (default)
python3 skills/video-understand/scripts/understand_video.py video.mp4

# Keyframe extraction
python3 skills/video-understand/scripts/understand_video.py video.mp4 -m keyframe

# Regular interval extraction
python3 skills/video-understand/scripts/understand_video.py video.mp4 -m interval

# Limit frames extracted
python3 skills/video-understand/scripts/understand_video.py video.mp4 --max-frames 10

# Use a larger Whisper model
python3 skills/video-understand/scripts/understand_video.py video.mp4 --whisper-model small

# Frames only, skip transcription
python3 skills/video-understand/scripts/understand_video.py video.mp4 --no-transcribe

# Quiet mode (JSON only, no progress)
python3 skills/video-understand/scripts/understand_video.py video.mp4 -q

# Output to file
python3 skills/video-understand/scripts/understand_video.py video.mp4 -o result.json

CLI Options

FlagDescription
videoInput video file (positional, required)
-m, --modeExtraction mode: scene (default), keyframe, interval
--max-framesMaximum frames to keep (default: 20)
--whisper-modelWhisper model size: tiny, base, small, medium, large (default: base)
--no-transcribeSkip audio transcription, extract frames only
-o, --outputWrite result JSON to file instead of stdout
-q, --quietSuppress progress messages, output only JSON

Extraction Modes

ModeHow it worksBest for
sceneDetects scene changes via ffmpeg select='gt(scene,0.3)'Most videos, varied content
keyframeExtracts I-frames (codec keyframes)Encoded video with natural keyframe placement
intervalEvenly spaced frames based on duration and max-framesFixed sampling, predictable output

If scene mode detects no scene changes, it automatically falls back to interval mode.

Output

The script outputs JSON to stdout (or file with -o). See references/output-format.md for the full schema.

json
{
  "video": "video.mp4",
  "duration": 18.076,
  "resolution": {"width": 1224, "height": 1080},
  "mode": "scene",
  "frames": [
    {"path": "/abs/path/frame_0001.jpg", "timestamp": 0.0, "timestamp_formatted": "00:00"}
  ],
  "frame_count": 12,
  "transcript": [
    {"start": 0.0, "end": 2.5, "text": "Hello and welcome..."}
  ],
  "text": "Full transcript...",
  "note": "Use the Read tool to view frame images for visual understanding."
}

Use the Read tool on frame image paths to visually inspect extracted frames.

References

  • references/output-format.md -- Full JSON output schema documentation

© calesthio, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts, references) in .agents/skills/video-understand of calesthio/OpenMontage.

  • SKILL.md
  • references/output-format.md
  • scripts/understand_video.py

Open the folder on GitHubat commit 9327439

Compare with similar skills

Video Understand next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Understand compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Understand this skillcalesthio/OpenMontage66k—~841Automated safety check: PassAGPL-3.0
Karaoke CaptionsAI-Builder-Club/skills1.3k—~850Automated safety check: PassNone
Ffmpegrendi-api/ffmpeg-cheatsheet1.7k—~1.2kAutomated safety check: PassNone
AutoshortsUpload-Post/skill-autoshorts151—~5.3kAutomated safety check: NotesMIT
Stage EditOrkas-AI/Orkas-VideoStudio499—~2.4kAutomated safety check: PassMIT
Proof Videoopenclaw/openclaw392k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Karaoke Captions

    AI-Builder-Club/skills

    Generate TikTok/Shorts-style karaoke captions using MLX Whisper, ASS subtitles, and FFmpeg libass.

    1.3k GitHub stars~850 tokensUpdated 25 days ago
    Media & CreativeAuto-check passed
  • Ffmpeg

    rendi-api/ffmpeg-cheatsheet

    A skill your agent uses when the user asks for FFmpeg or FFprobe commands, video/audio conversion, trimming, resizing, padding, overlays, subtitles, thumbnails, GIFs, storyboards, slideshows…

    1.7k GitHub stars~1.2k tokensUpdated 5 mo ago
    Media & CreativeAuto-check passed
  • Autoshorts

    Upload-Post/skill-autoshorts

    Daily pipeline that picks one long video from a folder, transcribes it with Whisper, uses Gemini 3 Flash multimodal to find every viral short-form moment, cuts each candidate with FFmpeg, adds a…

    151 GitHub stars~5.3k tokensUpdated 5 mo ago
    Media & CreativeAuto-check: notes
  • Stage Edit

    Orkas-AI/Orkas-VideoStudio

    Intelligent editing of real user-supplied footage—understand it with transcript/inspected-frame/scene/silence/quality evidence, then choose deterministic timeline operations or a constrained…

    499 GitHub stars~2.4k tokensUpdated 18 days ago
    Media & CreativeAuto-check passed
  • Proof Video

    openclaw/openclaw

    Add subtitles, captions, narration cues, or zoom to a proof video or PR recording using repo-local capture helpers and a system ffmpeg renderer.

    392k GitHub stars~2.4k tokensUpdated today
    Media & CreativeAuto-check passed
  • Video Production

    ffroliva/gflow-cli

    A skill your agent uses when the user wants a finished video out of gflow rather than a single clip — a scripted scene, a talking-head or dialogue piece, an explainer, a product montage, a story…

    269 GitHub stars~7k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed

More from calesthio/OpenMontage

All 41 skills in this repo
  • Avatar Video

    calesthio/OpenMontage

    Create AI avatar videos with precise control over avatars, voices, scripts, scenes, and backgrounds using HeyGen's v2 API.

    66k GitHub stars~1.6k tokensUpdated 7 days ago
    Auto-check passed
  • D3 Viz

    calesthio/OpenMontage

    Creating interactive data visualisations using d3.js. An agent skill from calesthio/OpenMontage.

    66k GitHub starsUsed in 3 repos~5.4k tokens
    Auto-check passed
  • Create Video

    calesthio/OpenMontage

    Create videos from a text prompt using HeyGen's Video Agent.

    66k GitHub stars~1.3k tokensUpdated 7 days ago
    Auto-check passed
  • Threejs World Generation

    calesthio/OpenMontage

    Build deterministic, editable, free-viewpoint Three.js worlds from text or structured briefs.

    66k GitHub stars~2k tokensUpdated 7 days ago
    Auto-check passed
  • Video Edit

    calesthio/OpenMontage

    Edit videos locally using ffmpeg. An agent skill from calesthio/OpenMontage.

    66k GitHub stars~855 tokensUpdated 7 days ago
    Auto-check: notes
  • Visual Style

    calesthio/OpenMontage

    Create, extract, and apply portable visual design systems via visual-style.md files.

    66k GitHub stars~1.5k tokensUpdated 7 days ago
    Auto-check passed

Works with

Questions about Video Understand

What does Video Understand do?

Understand video content locally using ffmpeg frame extraction and Whisper transcription. Video Understand is an agent skill from calesthio/OpenMontage. Understand video content locally using ffmpeg frame extraction and Whisper transcription.

When should I use Video Understand?

Video Understand fits situations like: understanding what a video contains; transcribing video audio locally; extracting key frames for visual analysis; getting video content without API keys.

How do I install Video Understand in Claude Code?

Run `npx skills add calesthio/OpenMontage --skill video-understand -a claude-code`. Or copy the skill folder (.agents/skills/video-understand in calesthio/OpenMontage) into .claude/skills/video-understand in your project. Claude Code loads it when a task matches its description.

How do I install Video Understand in Codex?

Run `npx skills add calesthio/OpenMontage --skill video-understand -a codex`. Or copy the skill folder (.agents/skills/video-understand in calesthio/OpenMontage) into .agents/skills/video-understand in your project. Codex loads it when a task matches its description.

Can I use Video Understand in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add calesthio/OpenMontage --skill video-understand -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-understand, .gemini/skills/video-understand, .github/skills/video-understand and .opencode/skills/video-understand in your project.

What does Video Understand need to run?

Going by SKILL.md and its folder, Video Understand needs Python for the scripts in its folder and the command-line tools its instructions call (python3, brew and pip). Our summary lists: Python 3.

Does Video Understand access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Video Understand safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Understand use?

Video Understand is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Video Understand use?

About 841 tokens (SKILL.md is roughly 3.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 784 tokens, read only when the agent opens those files.

What are the alternatives to Video Understand?

Skills that share tags, products or a category with Video Understand: Karaoke Captions (AI-Builder-Club/skills, 1.3k stars), Ffmpeg (rendi-api/ffmpeg-cheatsheet, 1.7k stars), Autoshorts (Upload-Post/skill-autoshorts, 151 stars) and Stage Edit (Orkas-AI/Orkas-VideoStudio, 499 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Understand?

calesthio (a GitHub user) maintains it in calesthio/OpenMontage, which has 65,930 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 3, 2026.

Source: calesthio/OpenMontage on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.