Agent skill

Gemini Video Understanding

by benchflow-ai in benchflow-ai/skillsbench

Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL…

Apache-2.0Auto-check passedMedia & Creative

Install Gemini Video Understanding

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill gemini-video-understanding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench gemini-video-understanding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-video-understanding .claude/skills/gemini-video-understanding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-video-understanding
GitHub stars
1.8k
Token cost
~2.4k tokens
SKILL.md length
536 words
Files
1
Skills in repo
189
Repo updated
First seen
Licence
Apache-2.0

At a glance

Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL…

  • Tasks that involve Computer vision
  • SKILL.md covers Purpose, When to Use, Required Libraries and Input Requirements, plus 7 more sections
  • Reaches youtube.com; needs GEMINI_API_KEY
  • Tasks that involve Transcription

What it does

Gemini Video Understanding is an agent skill from benchflow-ai/skillsbench. Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL analysis).

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Media & Creative, covering Computer vision and Transcription. It works with Google Gemini and YouTube. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Computer vision
  • Tasks that involve Transcription

Example prompts

  • “/gemini-video-understanding”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    Also links to:

    • ai.google.dev
    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Video Understanding loads about 2.4k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 536 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 536 words, ~2,396 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-video-understanding/SKILL.md (or your agent's skills folder).
name
gemini-video-understanding
description
Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL analysis).

Gemini Video Understanding Skill

Purpose

This skill enables video understanding workflows using the Google Gemini API, including video summarization, question answering, transcription with optional visual descriptions, timestamp-based queries (MM:SS), scene/timeline detection, video clipping, custom FPS sampling, multi-video comparison, and YouTube URL analysis.

When to Use

  • Summarizing a video into key points or chapters
  • Answering questions about what happens at specific timestamps (MM:SS)
  • Producing a transcript (optionally with visual context) and speaker labels
  • Detecting scene changes or building a timeline of events
  • Analyzing long videos by clipping to relevant segments or reducing FPS
  • Comparing multiple videos (up to 10 videos on Gemini 2.5+)
  • Analyzing public YouTube videos directly via URL

Required Libraries

The following Python libraries are required:

python
from google import genai
from google.genai import types
import os
import time

Input Requirements

  • File formats: MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP
  • Size constraints:
    • Use inline bytes for small files (rule of thumb: <20MB).
    • Use the File API upload flow for larger videos (most real videos).
  • YouTube:
    • Video must be public (not private/unlisted) and not age-restricted.
    • Provide a valid YouTube URL.
  • Duration / context window (model-dependent):
    • 2M-token models: ~2 hours (default resolution) or ~6 hours (low-res).
    • 1M-token models: ~1 hour (default) or ~3 hours (low-res).
  • Timestamps: Use MM:SS (e.g., 01:15) when requesting time-based answers.

Output Schema

All extracted/derived content should be returned as valid JSON conforming to this schema:

json
{
  "success": true,
  "source": {
    "type": "file|youtube",
    "id": "video.mp4|VIDEO_ID_OR_URL",
    "model": "gemini-2.5-flash"
  },
  "summary": "Concise summary of the video...",
  "transcript": {
    "available": true,
    "text": "Full transcript text (may include speaker labels)...",
    "includes_visual_descriptions": true
  },
  "events": [
    {
      "timestamp": "MM:SS",
      "description": "What happens at this time",
      "category": "scene_change|key_point|action|other"
    }
  ],
  "warnings": [
    "Optional warnings about limitations, missing timestamps, or low confidence areas"
  ]
}
Field Descriptions
  • success: Whether the analysis completed successfully
  • source.type: file for uploaded/local content, youtube for YouTube analysis
  • source.id: Filename for local uploads, or URL/ID for YouTube
  • source.model: Gemini model used for the request
  • summary: High-level video summary
  • transcript.*: Transcript payload (may be omitted or available=false if not requested)
  • events: Timeline items with MM:SS timestamps (chapters, scene changes, key actions)
  • warnings: Any issues that could affect correctness (e.g., “timestamp not found”, “long video clipped”)

Code Examples

Basic Video Analysis (Local Video + File API)
python
from google import genai
import os
import time

client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))

# Upload video (File API for >20MB)
myfile = client.files.upload(file="video.mp4")

# Wait for processing
while myfile.state.name == "PROCESSING":
    time.sleep(1)
    myfile = client.files.get(name=myfile.name)

if myfile.state.name == "FAILED":
    raise ValueError("Video processing failed")

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=["Summarize this video in 3 key points", myfile],
)

print(response.text)
YouTube Video Analysis (Public Videos Only)
python
from google import genai
from google.genai import types
import os

client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=[
        "Summarize the main topics discussed",
        types.Part.from_uri(
            uri="https://www.youtube.com/watch?v=VIDEO_ID",
            mime_type="video/mp4",
        ),
    ],
)

print(response.text)
Inline Video (<20MB)
python
from google import genai
from google.genai import types
import os

client = genai.Client(api_key=os.getenv("GEMINI_API_KEY"))

with open("short-clip.mp4", "rb") as f:
    video_bytes = f.read()

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=[
        "What happens in this video?",
        types.Part.from_bytes(data=video_bytes, mime_type="video/mp4"),
    ],
)

print(response.text)
Video Clipping (Analyze a Segment Only)
python
from google.genai import types

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=[
        "Summarize this segment",
        types.Part.from_video_metadata(
            file_uri=myfile.uri,
            start_offset="40s",
            end_offset="80s",
        ),
    ],
)
Custom Frame Rate (Token/Cost Control)
python
from google.genai import types

# Lower FPS for static content (saves tokens)
response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=[
        "Analyze this presentation",
        types.Part.from_video_metadata(file_uri=myfile.uri, fps=0.5),
    ],
)

# Higher FPS for fast-moving content
response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=[
        "Analyze rapid movements in this sports video",
        types.Part.from_video_metadata(file_uri=myfile.uri, fps=5),
    ],
)
Timeline / Scene Detection (MM:SS)
python
response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=[
        """Create a timeline with timestamps:
        - Key events
        - Scene changes
        - Important moments
        Format: MM:SS - Description
        """,
        myfile,
    ],
)
Show full SKILL.md (214 more words)Show less
Transcription (Optional Visual Descriptions)
python
response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=[
        """Transcribe with visual context:
        - Audio transcription
        - Visual descriptions of important moments
        - Timestamps for salient events
        """,
        myfile,
    ],
)
Structured JSON Output (Schema-Guided)
python
from pydantic import BaseModel
from typing import List
from google.genai import types as genai_types

class VideoEvent(BaseModel):
    timestamp: str  # MM:SS
    description: str
    category: str

class VideoAnalysis(BaseModel):
    summary: str
    events: List[VideoEvent]
    duration: str

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=["Analyze this video", myfile],
    config=genai_types.GenerateContentConfig(
        response_mime_type="application/json",
        response_schema=VideoAnalysis,
    ),
)

Best Practices

  • Use the File API for most videos (>20MB) and wait for processing to complete before analysis.
  • Reduce token usage by clipping to the relevant segment and/or lowering FPS for static content.
  • Improve accuracy by being explicit about the desired output (timestamps format, number of events, whether you want scene changes vs actions vs chapters).
  • Use gemini-2.5-pro when you need highest-quality reasoning over complex, long, or visually dense videos.

Error Handling

python
import time

def upload_and_wait(client, file_path: str, max_wait_s: int = 300):
    myfile = client.files.upload(file=file_path)
    waited = 0

    while myfile.state.name == "PROCESSING" and waited < max_wait_s:
        time.sleep(5)
        waited += 5
        myfile = client.files.get(name=myfile.name)

    if myfile.state.name == "FAILED":
        raise ValueError(f"Video processing failed: {myfile.state.name}")
    if myfile.state.name == "PROCESSING":
        raise TimeoutError(f"Processing timeout after {max_wait_s}s")

    return myfile

Common issues:

  • Upload processing stuck: wait and poll; fail after a max timeout.
  • YouTube errors: verify the video is public and not age-restricted.
  • Rate limits: retry with exponential backoff.
  • Incorrect timestamps: re-prompt with strict “MM:SS” formatting and request fewer events.

Limitations

  • Long-video support is limited by model context and token budget (default vs low-res modes).
  • YouTube analysis requires public videos; live streaming analysis is not supported.
  • Very long videos may require chunking (clip by time range and process in segments).
  • Multi-video comparison is limited (up to 10 videos per request on Gemini 2.5+).

Version History

  • 1.0.0 (2026-01-15): Initial release focused on Gemini video understanding (summaries, Q&A, timestamps, clipping, FPS control, YouTube, and structured outputs).

Resources

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-video-understanding of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Gemini Video Understanding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Video Understanding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Video Understanding this skillbenchflow-ai/skillsbench1.8k—~2.4kAutomated safety check: PassApache-2.0
Gemini Video Understandingeinverne/dotfiles121—~2.6kAutomated safety check: NotesMIT
AI MultimodalMicrock/ordinary-claude-skills404—~2.7kAutomated safety check: NotesMIT
Watch Videocoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Gemini Yt Video Transcriptsundial-org/awesome-openclaw-skills663—~293Automated safety check: PassNone
Watching Videosoxbshw/watch-skill470—~599Automated safety check: NotesMIT

Similar skills

  • Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

    121 GitHub stars~2.6k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check: notes
  • AI Multimodal

    Microck/ordinary-claude-skills

    Process and generate multimedia content using Google Gemini API.

    404 GitHub stars~2.7k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    851 GitHub stars~3.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Gemini Yt Video Transcript

    sundial-org/awesome-openclaw-skills

    Create a verbatim transcript for a YouTube URL using Google Gemini (speaker labels, paragraph breaks; no time codes).

    663 GitHub stars~293 tokensUpdated 7 mo ago
    Media & CreativeAuto-check passed
  • Watching Videos

    oxbshw/watch-skill

    The user shared a video URL, a YouTube/TikTok/stream link, a local video file, a screen recording, a meeting recording, or a playlist/folder of videos — "watch this", "summarize this video", "what's…

    470 GitHub stars~599 tokensUpdated 25 days ago
    Media & CreativeAuto-check: notes
  • 9Router Speech-to-Text

    decolua/9router

    Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others.

    31k GitHub stars~914 tokensUpdated 2 days ago
    Media & CreativeAuto-check passed

More from benchflow-ai/skillsbench

All 189 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Gemini Video Understanding

What does Gemini Video Understanding do?

Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL…. Gemini Video Understanding is an agent skill from benchflow-ai/skillsbench. Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL analysis).

When should I use Gemini Video Understanding?

Gemini Video Understanding fits situations like: tasks that involve Computer vision; tasks that involve Transcription.

How do I install Gemini Video Understanding in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill gemini-video-understanding -a claude-code`. Or copy the skill folder (tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-video-understanding in benchflow-ai/skillsbench) into .claude/skills/gemini-video-understanding in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Video Understanding in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill gemini-video-understanding -a codex`. Or copy the skill folder (tasks-extra/pedestrian-traffic-counting/environment/skills/gemini-video-understanding in benchflow-ai/skillsbench) into .agents/skills/gemini-video-understanding in your project. Codex loads it when a task matches its description.

Can I use Gemini Video Understanding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill gemini-video-understanding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-video-understanding, .gemini/skills/gemini-video-understanding, .github/skills/gemini-video-understanding and .opencode/skills/gemini-video-understanding in your project.

What does Gemini Video Understanding need to run?

Going by SKILL.md and its folder, Gemini Video Understanding needs credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY.

Does Gemini Video Understanding access the network?

SKILL.md names 3 domains. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. As links in the text: ai.google.dev and aistudio.google.com. This is read from the text; nothing was executed.

Is Gemini Video Understanding safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Gemini Video Understanding use?

Gemini Video Understanding is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Video Understanding use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemini Video Understanding?

Skills that share tags, products or a category with Gemini Video Understanding: Gemini Video Understanding (einverne/dotfiles, 121 stars), AI Multimodal (Microck/ordinary-claude-skills, 404 stars), Watch Video (coreyhaines31/makerskills, 851 stars) and Gemini Yt Video Transcript (sundial-org/awesome-openclaw-skills, 663 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Video Understanding?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,835 GitHub stars. The repository holds 189 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.