Agent skill

Gemini Video Understanding

by einverne in einverne/dotfiles

Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

MITAuto-check: notesAI & LLM Engineering

Install Gemini Video Understanding

skills CLI
$ npx skills add einverne/dotfiles --skill gemini-video-understanding -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install einverne/dotfiles gemini-video-understanding --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/einverne/dotfiles.git skills-src && mkdir -p .claude/skills && cp -r skills-src/claude/skills/gemini-video-understanding .claude/skills/gemini-video-understanding && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gemini-video-understanding
GitHub stars
121
Token cost
~2.6k tokens
SKILL.md length
719 words
Files
8 (incl. scripts)
Skills in repo
39
Repo updated
First seen
Licence
MIT

At a glance

Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

  • Works in 6 steps: Video Summarization → Educational Content → Timestamp-Specific Questions → …
  • Tasks that involve Transcription
  • SKILL.md covers Capabilities, Supported Formats, Models Available and API Key Configuration, plus 9 more sections
  • Runs Python scripts from its folder; calls python; reaches youtube.com; needs GEMINI_API_KEY

What it does

Gemini Video Understanding is an agent skill from einverne/dotfiles. Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs. Supports 9 video formats, multiple models (Gemini 2.5/2.0), and context windows up to 2M tokens (6 hours of video).

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts (for example `EXAMPLES.md`, `QUICKSTART.md` and `README.md`).

It sits in AI & LLM Engineering, covering Transcription, Computer vision and LLM API integration. It works with Google Gemini and YouTube. The repository describes itself as: my personal dotfiles managed by dotbot, zinit. The licence is MIT.

When your agent uses it

  • Tasks that involve Transcription
  • Tasks that involve Computer vision
  • Tasks that involve LLM API integration

Example prompts

  • “/gemini-video-understanding”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read, Write

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Video Summarization
  2. Educational Content
  3. Timestamp-Specific Questions
  4. Transcription
  5. Content Comparison
  6. Action Detection

What it can do on your machine

Read from SKILL.md and the folder at commit c6c0686. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • youtube.com

    Also links to:

    • ai.google.dev
    • aistudio.google.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Gemini Video Understanding loads about 2.6k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 719 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:54
    claude/skills/gemini-video-understanding/.env`
  • NoteMentions a .env fileSKILL.md:55
    3. **Project root**: `.env` file in project root
  • NoteMentions a .env fileSKILL.md:63
    # Option 2: Skill directory .env file
  • NoteMentions a .env fileSKILL.md:64
    claude/skills/gemini-video-understanding/.env
  • NoteMentions a .env fileSKILL.md:66
    # Option 3: Project root .env file
  • NoteMentions a .env fileSKILL.md:67
    cho "GEMINI_API_KEY=your-api-key-here" > .env
  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from einverne/dotfiles at commit c6c0686, republished under its MIT licence (© einverne). 719 words, ~2,650 tokens.

Download SKILL.mdSave it as .claude/skills/gemini-video-understanding/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
gemini-video-understanding
description
Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs. Supports 9 video formats, multiple models (Gemini 2.5/2.0), and context windows up to 2M tokens (6 hours of video).
allowed-tools
Bash, Read, Write
license
MIT
metadata.version
1.0.0
metadata.author
ClaudeKit
metadata.api-provider
Google Gemini
metadata.requires-api-key
GEMINI_API_KEY

Gemini Video Understanding Skill

This skill enables comprehensive video analysis using Google's Gemini API, including video summarization, question answering, transcription, timestamp references, and more.

Capabilities

  • Video Summarization: Create concise summaries of video content
  • Question Answering: Answer specific questions about video content
  • Transcription: Transcribe audio with visual descriptions and timestamps
  • Timestamp References: Query specific moments in videos (MM:SS format)
  • Video Clipping: Process specific segments using start/end offsets
  • Multiple Videos: Compare and analyze up to 10 videos (Gemini 2.5+)
  • YouTube Support: Analyze YouTube videos directly (preview feature)
  • Custom Frame Rate: Adjust FPS sampling for different video types

Supported Formats

  • MP4, MPEG, MOV, AVI, FLV, MPG, WebM, WMV, 3GPP

Models Available

Gemini 2.5 Series:

  • gemini-2.5-pro - Best quality, 1M context
  • gemini-2.5-flash - Balanced quality/speed, 1M context
  • gemini-2.5-flash-preview-09-2025 - Preview features, 1M context

Gemini 2.0 Series:

  • gemini-2.0-flash - Fast processing
  • gemini-2.0-flash-lite - Lightweight option

Context Windows:

  • 2M token models: ~2 hours (default) or ~6 hours (low-res)
  • 1M token models: ~1 hour (default) or ~3 hours (low-res)

API Key Configuration

The skill checks for GEMINI_API_KEY in this order:

  1. Process environment: process.env.GEMINI_API_KEY or $GEMINI_API_KEY
  2. Skill directory: .claude/skills/gemini-video-understanding/.env
  3. Project root: .env file in project root

To set up:

bash
# Option 1: Environment variable (recommended)
export GEMINI_API_KEY="your-api-key-here"

# Option 2: Skill directory .env file
echo "GEMINI_API_KEY=your-api-key-here" > .claude/skills/gemini-video-understanding/.env

# Option 3: Project root .env file
echo "GEMINI_API_KEY=your-api-key-here" > .env

Get your API key at: https://aistudio.google.com/apikey

Usage Instructions

When to Use This Skill

Use this skill when the user asks to:

  • Analyze, summarize, or describe video content
  • Answer questions about videos
  • Transcribe video audio with visual context
  • Extract information from specific timestamps
  • Compare multiple videos
  • Process YouTube video content
  • Create quizzes or educational content from videos
Basic Video Analysis

For video files:

python
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-path "/path/to/video.mp4" \
  --prompt "Summarize this video in 3 key points"

For YouTube URLs:

python
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --youtube-url "https://www.youtube.com/watch?v=VIDEO_ID" \
  --prompt "What are the main topics discussed?"
Advanced Features

Video Clipping (specific time range):

python
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-path "/path/to/video.mp4" \
  --prompt "Summarize this segment" \
  --start-offset "40s" \
  --end-offset "80s"

Custom Frame Rate:

python
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-path "/path/to/video.mp4" \
  --prompt "Analyze the rapid movements" \
  --fps 5

Transcription with Timestamps:

python
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-path "/path/to/video.mp4" \
  --prompt "Transcribe the audio with timestamps and visual descriptions"

Multiple Videos (Gemini 2.5+ only):

python
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-paths "/path/video1.mp4" "/path/video2.mp4" \
  --prompt "Compare these two videos and highlight the differences"

Model Selection:

python
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-path "/path/to/video.mp4" \
  --prompt "Detailed analysis" \
  --model "gemini-2.5-pro"
Script Parameters
Required (one of):
  --video-path PATH           Path to local video file
  --youtube-url URL           YouTube video URL
  --video-paths PATH [PATH..] Multiple video paths (Gemini 2.5+)

Required:
  --prompt TEXT              Analysis prompt/question

Optional:
  --model NAME               Model to use (default: gemini-2.5-flash)
  --start-offset TIME        Video clip start (e.g., "40s", "1m30s")
  --end-offset TIME          Video clip end (e.g., "80s", "2m")
  --fps NUMBER               Frame sampling rate (default: 1)
  --output-file PATH         Save response to file
  --verbose                  Show detailed processing info

Common Use Cases

1. Video Summarization
Prompt: "Summarize this video in 3 key points with timestamps"
2. Educational Content
Prompt: "Create a quiz with 5 questions and answer key based on this video"
3. Timestamp-Specific Questions
Prompt: "What happens at 01:15 and how does it relate to the topic at 02:30?"
4. Transcription
Prompt: "Transcribe the audio from this video with timestamps for salient events and visual descriptions"
5. Content Comparison
Prompt: "Compare these two product demo videos. Which one explains the features more clearly?"
6. Action Detection
Prompt: "List all the actions performed in this tutorial video with timestamps"

Rate Limits & Quotas

Free Tier (per model):

  • 10-15 RPM (requests per minute)
  • 1M-4M TPM (tokens per minute)
  • 1,500 RPD (requests per day)

YouTube Limitations:

  • Free tier: 8 hours of YouTube video per day
  • Paid tier: No length-based limits
  • Public videos only (no private/unlisted)

Storage (Files API):

  • 20GB per project
  • 2GB per file
  • 48-hour retention period

Token Calculation

Video tokens depend on resolution:

  • Default resolution: ~300 tokens per second of video
  • Low resolution: ~100 tokens per second of video

Example: A 10-minute video = 600 seconds × 300 tokens = ~180,000 tokens

Error Handling

Common errors and solutions:

ErrorCauseSolution
400 Bad RequestInvalid video format or corrupt fileCheck file format and integrity
403 ForbiddenInvalid/missing API keyVerify GEMINI_API_KEY configuration
404 Not FoundFile URI not foundEnsure file is uploaded and active
429 Too Many RequestsRate limit exceededImplement backoff, upgrade to paid tier
500 Internal ErrorServer-side issueRetry with exponential backoff
Show full SKILL.md (269 more words)Show less

Best Practices

  1. Use Files API for videos >20MB - More reliable than inline data
  2. Wait for file processing - Poll until state is ACTIVE before analysis
  3. Optimize FPS - Use lower FPS for static content to save tokens
  4. Clip long videos - Process specific segments instead of entire video
  5. Cache context - Reuse uploaded files for multiple queries
  6. Batch processing - Process multiple short videos in one request (2.5+)
  7. Specific prompts - Be precise about what you want to extract

Implementation Notes

For Claude Code:

When a user requests video analysis:

  1. Check API key availability first using the helper script
  2. Determine video source: local file, YouTube URL, or multiple videos
  3. Select appropriate model based on requirements (default: gemini-2.5-flash)
  4. Run the analysis script with proper parameters
  5. Parse and present results to the user clearly
  6. Handle errors gracefully with helpful suggestions
Files API Workflow:

For videos >20MB or reusable content:

  1. Upload video using Files API (script handles this automatically)
  2. Wait for ACTIVE state (polling included in script)
  3. Use file URI for analysis
  4. Files auto-delete after 48 hours
Inline Data Workflow:

For videos <20MB:

  1. Read video file as bytes
  2. Base64 encode for API
  3. Send in generateContent request
  4. Single-use, no upload needed

Example Workflows

Workflow 1: YouTube Video Summary
bash
# User: "Analyze this YouTube tutorial video"
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --youtube-url "https://www.youtube.com/watch?v=abc123" \
  --prompt "Create a structured summary with: 1) Main topics, 2) Key takeaways, 3) Recommended audience"
Workflow 2: Interview Transcription
bash
# User: "Transcribe this interview with timestamps"
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-path "interview.mp4" \
  --prompt "Transcribe this interview with speaker labels, timestamps, and visual descriptions of gestures or slides shown"
Workflow 3: Product Comparison
bash
# User: "Compare these two product demo videos"
python .claude/skills/gemini-video-understanding/scripts/analyze_video.py \
  --video-paths "demo1.mp4" "demo2.mp4" \
  --model "gemini-2.5-pro" \
  --prompt "Compare these product demos on: features shown, presentation quality, clarity of explanation, and overall effectiveness"

Troubleshooting

API Key Not Found:

bash
# Check API key detection
python .claude/skills/gemini-video-understanding/scripts/check_api_key.py

Video Too Large:

Error: Request size exceeds 20MB
Solution: Script automatically uses Files API for large videos

Processing Timeout:

Error: File not reaching ACTIVE state
Solution: Check video integrity, try smaller file, or different format

Rate Limit Errors:

Error: 429 Too Many Requests
Solution: Wait before retry, or upgrade to paid tier

Additional Resources

Version History

  • 1.0.0 (2025-10-26): Initial release with full video understanding capabilities

© einverne, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts) in claude/skills/gemini-video-understanding of einverne/dotfiles.

  • SKILL.md
  • .env.example
  • EXAMPLES.md
  • QUICKSTART.md
  • README.md
  • requirements.txt
  • scripts/analyze_video.py
  • scripts/check_api_key.py

Open the folder on GitHubat commit c6c0686

Compare with similar skills

Gemini Video Understanding next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Gemini Video Understanding compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Gemini Video Understanding this skilleinverne/dotfiles121—~2.6kAutomated safety check: NotesMIT
AI MultimodalMicrock/ordinary-claude-skills404—~2.7kAutomated safety check: NotesMIT
Watch Videocoreyhaines31/makerskills851—~3.8kAutomated safety check: PassMIT
Gemini Video Understandingbenchflow-ai/skillsbench1.8k—~2.4kAutomated safety check: PassApache-2.0
ModLens Image Vision Bridgeliustack/modlens4.2k—~1.3kAutomated safety check: NotesMIT
Gemini Live API Devgoogle-gemini/gemini-skills4.3k—~4.6kAutomated safety check: PassApache-2.0

Similar skills

  • AI Multimodal

    Microck/ordinary-claude-skills

    Process and generate multimedia content using Google Gemini API.

    404 GitHub stars~2.7k tokensUpdated 1 mo ago
    Media & CreativeAuto-check: notes
  • Watch Video

    coreyhaines31/makerskills

    When you want to extract content from a video — YouTube, Loom, Vimeo, Riverside, Zoom recording, local MP4, X/IG video, anything yt-dlp supports.

    851 GitHub stars~3.8k tokensUpdated 2 days ago
    Media & CreativeAuto-check passed
  • Gemini Video Understanding

    benchflow-ai/skillsbench

    Analyze videos with Google Gemini API (summaries, Q&A, transcription with timestamps + visual context, scene/timeline detection, video clipping, FPS control, multi-video comparison, and YouTube URL…

    1.8k GitHub stars~2.4k tokensUpdated 2 mo ago
    Media & CreativeAuto-check passed
  • Gives text-only models sight by running the modlens CLI on an image path or URL and returning structured JSON evidence with transcribed text, layout and semantics.

    4.2k GitHub stars~1.3k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check: notes
  • Gemini Live API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when building real-time, bidirectional streaming applications with the Gemini Live API, or migrating legacy Live models (2.0/2.5/3.1) to Gemini 3.8 Live.

    4.3k GitHub stars~4.6k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed
  • Summarize

    swarmclawai/swarmclaw

    Summarize or extract text/transcripts from URLs, podcasts, YouTube videos, and local files using the summarize CLI.

    689 GitHub stars~531 tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed

More from einverne/dotfiles

All 39 skills in this repo
  • DOCX

    einverne/dotfiles

    Comprehensive document creation, editing, and analysis with support for tracked changes, comments, formatting preservation, and text extraction.

    121 GitHub starsUsed in 35 repos~2.5k tokens
    Auto-check: notes
  • Chrome Devtools

    einverne/dotfiles

    Browser automation, debugging, and performance analysis using Puppeteer CLI scripts.

    121 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • PDF

    einverne/dotfiles

    Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms.

    121 GitHub starsUsed in 47 repos~1.8k tokens
    Auto-check passed
  • Gemini Audio

    einverne/dotfiles

    Guide for implementing Google Gemini API audio capabilities - analyze audio with transcription, summarization, and understanding (up to 9.5 hours), plus generate speech with controllable TTS.

    121 GitHub stars~2k tokensUpdated 1 mo ago
    Auto-check: notes
  • Guide for implementing Google Gemini API document processing - analyze PDFs with native vision to extract text, images, diagrams, charts, and tables.

    121 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check: notes
  • PPTX

    einverne/dotfiles

    Presentation creation, editing, and analysis. An agent skill from einverne/dotfiles.

    121 GitHub starsUsed in 38 repos~6.4k tokens
    Auto-check: notes

Questions about Gemini Video Understanding

What does Gemini Video Understanding do?

Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs. Gemini Video Understanding is an agent skill from einverne/dotfiles. Analyze videos using Google's Gemini API - describe content, answer questions, transcribe audio with visual descriptions, reference timestamps, clip videos, and process YouTube URLs.

When should I use Gemini Video Understanding?

Gemini Video Understanding fits situations like: tasks that involve Transcription; tasks that involve Computer vision; tasks that involve LLM API integration.

How do I install Gemini Video Understanding in Claude Code?

Run `npx skills add einverne/dotfiles --skill gemini-video-understanding -a claude-code`. Or copy the skill folder (claude/skills/gemini-video-understanding in einverne/dotfiles) into .claude/skills/gemini-video-understanding in your project. Claude Code loads it when a task matches its description.

How do I install Gemini Video Understanding in Codex?

Run `npx skills add einverne/dotfiles --skill gemini-video-understanding -a codex`. Or copy the skill folder (claude/skills/gemini-video-understanding in einverne/dotfiles) into .agents/skills/gemini-video-understanding in your project. Codex loads it when a task matches its description.

Can I use Gemini Video Understanding in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add einverne/dotfiles --skill gemini-video-understanding -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gemini-video-understanding, .gemini/skills/gemini-video-understanding, .github/skills/gemini-video-understanding and .opencode/skills/gemini-video-understanding in your project.

What does Gemini Video Understanding need to run?

Going by SKILL.md and its folder, Gemini Video Understanding needs Python for the scripts in its folder, the command-line tools its instructions call (python) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read, Write.

Does Gemini Video Understanding access the network?

SKILL.md names 3 domains. In commands or code: youtube.com; the agent is likely to contact it when it follows the instructions. As links in the text: ai.google.dev and aistudio.google.com. This is read from the text; nothing was executed.

Is Gemini Video Understanding safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file; pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Gemini Video Understanding use?

Gemini Video Understanding is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Gemini Video Understanding use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Gemini Video Understanding?

Skills that share tags, products or a category with Gemini Video Understanding: AI Multimodal (Microck/ordinary-claude-skills, 404 stars), Watch Video (coreyhaines31/makerskills, 851 stars), Gemini Video Understanding (benchflow-ai/skillsbench, 1.8k stars) and ModLens Image Vision Bridge (liustack/modlens, 4.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Gemini Video Understanding?

einverne (a GitHub user) maintains it in einverne/dotfiles, which has 121 GitHub stars. The repository holds 39 skills in this directory. The repository was last updated on September 9, 2026.

Source: einverne/dotfiles on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.