Agent skill

Video Analyzer

by mikefutia in mikefutia/claude-vision

Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details…

No licenceAuto-check: notesMedia & Creative

Install Video Analyzer

skills CLI
$ npx skills add mikefutia/claude-vision --skill video-analyzer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mikefutia/claude-vision video-analyzer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
video-analyzer
GitHub stars
101
Token cost
~747 tokens
SKILL.md length
326 words
Files
4 (incl. scripts)
Skills in repo
1
Repo updated
First seen
Licence
None found

At a glance

Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details…

  • Works in 6 steps: Parse the arguments from $ARGUMENTS → Verify the video file exists at the… → Run the analysis script using the… → …
  • You need to understand what actually happens in a video
  • SKILL.md covers Prerequisites, Steps and Output
  • Runs Python scripts from its folder; calls python3; needs GEMINI_API_KEY

What it does

Video Analyzer is an agent skill from mikefutia/claude-vision. Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details, and key timestamped moments. Strong anti-hallucination guardrails — will not invent narrators, voiceovers, or speaker names. Use when you need to understand what actually happens in a video.

Its SKILL.md is about 750 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `README.md` and `scripts/analyze_video.py`).

It sits in Media & Creative, covering Text to speech and voice. It works with Google Gemini and Python. The repository describes itself as: Claude Vision Skill (Mike Futia | SCALE AI).

When your agent uses it

  • You need to understand what actually happens in a video
  • Tasks that involve Text to speech and voice

Example prompts

  • “silent”
  • “Use the video-analyzer skill to analyz a video file with Google Gemini and returns a structured markdown report covering top-level summary…”
  • “/video-analyzer”

Requirements

  • Python 3
  • A credential in GEMINI_API_KEY
  • Pre-approved tools (allowed-tools): Bash, Read

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Parse the arguments from $ARGUMENTS
  2. Verify the video file exists at the given path. If not, report the error and stop.
  3. Run the analysis script using the absolute path to its install location
  4. The script will
  5. Capture stdout and present the report to the user.
  6. If the script exits with an error, help the user troubleshoot

What it can do on your machine

Read from SKILL.md and the folder at commit 0665eb3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Read

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Video Analyzer loads about 747 tokens when it runs. Until then it costs about 100 tokens; SKILL.md has 326 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~100
When it runs · the whole SKILL.md, loaded when a task matches
~747

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Read

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 326 words (~747 tokens).

“Analyze a video file with Gemini and return a structured markdown report.”

— opening of SKILL.md by mikefutia
name
video-analyzer
allowed-tools
Bash, Read
argument-hint
<path/to/video.mp4> [--prompt "..."] [--fps N] [--model ...]
disable-model-invocation
true

Read the full SKILL.md on GitHub

Files

SKILL.md and 3 other files (scripts) in the repository root of mikefutia/claude-vision.

  • SKILL.md
  • .gitignore
  • README.md
  • scripts/analyze_video.py

Open the folder on GitHubat commit 0665eb3

Compare with similar skills

Video Analyzer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Video Analyzer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Video Analyzer this skillmikefutia/claude-vision101—~747Automated safety check: NotesNone
Gemini API Devgoogle-gemini/gemini-skills4.3k—~5.1kAutomated safety check: PassApache-2.0
Floe GuardFloe-Labs/floe-guard357—~853Automated safety check: PassMIT
Gemini Interactions APIAyuilos/Miffan217—~4.6kAutomated safety check: PassAGPL-3.0
Arkcli Code Examplevolcengine/ark-cli140—~743Automated safety check: PassApache-2.0
Watch Video Q&Abradautomates/claude-video18k—~4.3kAutomated safety check: NotesMIT

Similar skills

  • Gemini API Dev

    google-gemini/gemini-skills

    Official

    A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, speech generation (TTS), voice…

    4.3k GitHub stars~5.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Floe Guard

    Floe-Labs/floe-guard

    Know what every AI call really costs — floe-guard meters STT + TTS + LLM + telephony per call (Pipecat, LiveKit — Python & TypeScript), keeps a live ledger of real spend, and hard-stops the next…

    357 GitHub stars~853 tokensUpdated 5 days ago
    Media & CreativeAuto-check passed
  • A skill your agent uses when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses…

    217 GitHub stars~4.6k tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Arkcli Code Example

    volcengine/ark-cli

    arkcli +code-example:为指定基础模型生成多语言(Python / Go / Java / Node / curl)调用示例代码并写入本地文件。数据源是火山方舟 OpenTOP OpenGetSampleCode。当用户需要拿某个基础模型的 SDK / curl 调用示例、保存为本地接入模板时使用。反触发:TTS/ASR/语音模型没有 arkcli…

    140 GitHub stars~743 tokensUpdated yesterday
    Media & CreativeAuto-check passed
  • Watch Video Q&A

    bradautomates/claude-video

    Lets the agent answer questions about a video from a URL or local file by downloading it, extracting frames and a transcript, or by sending it to Gemini's video model.

    18k GitHub stars~4.3k tokensUpdated 15 days ago
    Media & CreativeAuto-check: notes
  • Vox Director

    Alisa0808/vox-director

    Turn ONE topic into a finished Vox-style paper-collage explainer / ad video, end to end on the Atlas Cloud API + local ffmpeg — script, collage keyframes, motion, voice-over, music, captions, all…

    2.2k GitHub stars~5.6k tokensUpdated 4 days ago
    Media & CreativeAuto-check passed

Questions about Video Analyzer

What does Video Analyzer do?

Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details…. Video Analyzer is an agent skill from mikefutia/claude-vision. Analyzes a video file with Google Gemini and returns a structured markdown report covering top-level summary, scene-by-scene breakdown, audio transcript (or honest "silent" note), visual details, and key timestamped moments.

When should I use Video Analyzer?

Video Analyzer fits situations like: you need to understand what actually happens in a video; tasks that involve Text to speech and voice.

How do I install Video Analyzer in Claude Code?

Run `npx skills add mikefutia/claude-vision --skill video-analyzer -a claude-code`. Or copy the skill folder (the mikefutia/claude-vision repository) into .claude/skills/video-analyzer in your project. Claude Code loads it when a task matches its description.

How do I install Video Analyzer in Codex?

Run `npx skills add mikefutia/claude-vision --skill video-analyzer -a codex`. Or copy the skill folder (the mikefutia/claude-vision repository) into .agents/skills/video-analyzer in your project. Codex loads it when a task matches its description.

Can I use Video Analyzer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mikefutia/claude-vision --skill video-analyzer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/video-analyzer, .gemini/skills/video-analyzer, .github/skills/video-analyzer and .opencode/skills/video-analyzer in your project.

What does Video Analyzer need to run?

Going by SKILL.md and its folder, Video Analyzer needs Python for the scripts in its folder, the command-line tools its instructions call (python3) and credentials named GEMINI_API_KEY. Our summary lists: Python 3; A credential in GEMINI_API_KEY. Its frontmatter pre-approves these tools: Bash, Read.

Does Video Analyzer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Video Analyzer safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Video Analyzer use?

No licence was found for Video Analyzer or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Video Analyzer use?

About 747 tokens (SKILL.md is roughly 3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Video Analyzer?

Skills that share tags, products or a category with Video Analyzer: Gemini API Dev (google-gemini/gemini-skills, 4.3k stars), Floe Guard (Floe-Labs/floe-guard, 357 stars), Gemini Interactions API (Ayuilos/Miffan, 217 stars) and Arkcli Code Example (volcengine/ark-cli, 140 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Video Analyzer?

mikefutia (a GitHub user) maintains it in mikefutia/claude-vision, which has 101 GitHub stars. The repository was last updated on May 4, 2026.

Source: mikefutia/claude-vision on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.